David Donofrio

dblp:45/7583 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16Software engineering, systems software and programming languages · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Storage systems · 34% Electronic design automation · 19% Memory systems · 19%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › domain-specific accelerator
data processing accelerator
0.412020
DRAM-Less: Hardware Acceleration of Data Processing with New Memory · HPCA 2020
Electronic design automation
hardware verification and test
0.412020
Efficiently Exploiting Low Activity Factors to Accelerate RTL Simulation · DAC 2020
Memory systems
non-volatile memory
0.412020
DRAM-Less: Hardware Acceleration of Data Processing with New Memory · HPCA 2020
Memory systems › non-volatile memory
phase change memory
0.412020
DRAM-Less: Hardware Acceleration of Data Processing with New Memory · HPCA 2020
Electronic design automation › hardware verification and test › hardware verification
RTL simulation
0.412020
Efficiently Exploiting Low Activity Factors to Accelerate RTL Simulation · DAC 2020
Storage systems › storage reliability
erasure coding
0.412019
Exploring Fault-Tolerant Erasure Codes for Scalable All-Flash Array Clusters · IEEE Trans. Parallel Distributed Syst. 2019
Storage systems
flash and SSD
0.412019
Exploring Fault-Tolerant Erasure Codes for Scalable All-Flash Array Clusters · IEEE Trans. Parallel Distributed Syst. 2019
Storage systems › data representation › data encoding › error correction coding
reed-solomon codes
0.412019
Exploring Fault-Tolerant Erasure Codes for Scalable All-Flash Array Clusters · IEEE Trans. Parallel Distributed Syst. 2019
Storage systems › flash and SSD › flash memory
flash storage
0.212016
NANDFlashSim: High-Fidelity, Microarchitecture-Aware NAND Flash Memory Simulation · ACM Trans. Storage 2016
Performance modeling and evaluation
simulation
0.212016
NANDFlashSim: High-Fidelity, Microarchitecture-Aware NAND Flash Memory Simulation · ACM Trans. Storage 2016
Performance modeling and evaluation › simulation › architectural simulation
simulation acceleration
0.112020
Efficiently Exploiting Low Activity Factors to Accelerate RTL Simulation · DAC 2020
Energy-efficient computing
energy-efficient architecture
0.112011
Hardware/software co-design for energy-efficient seismic modeling · SC 2011
Storage systems › file systems
distributed file system
0.112019
Exploring Fault-Tolerant Erasure Codes for Scalable All-Flash Array Clusters · IEEE Trans. Parallel Distributed Syst. 2019
Storage systems › file systems › distributed file system
parallel file system
0.112019
Exploring Fault-Tolerant Erasure Codes for Scalable All-Flash Array Clusters · IEEE Trans. Parallel Distributed Syst. 2019
Performance modeling and evaluation › benchmarking
storage benchmarking
0.112019
Exploring Fault-Tolerant Erasure Codes for Scalable All-Flash Array Clusters · IEEE Trans. Parallel Distributed Syst. 2019
Performance modeling and evaluation
workload characterization
0.112016
NANDFlashSim: High-Fidelity, Microarchitecture-Aware NAND Flash Memory Simulation · ACM Trans. Storage 2016
High-performance computing
scientific computing systems
0.012011
Hardware/software co-design for energy-efficient seismic modeling · SC 2011
High-performance computing › scientific computing systems
seismic imaging
0.012011
Hardware/software co-design for energy-efficient seismic modeling · SC 2011

Methods — techniques the papers use, named apart from their topics

signal reuse detection · 0.4acyclic partitioning algorithm · 0.4PRAM integration · 0.4PCIe accelerator emulation · 0.4replication · 0.4reed-solomon coding · 0.4timing model · 0.2microarchitecture simulation · 0.2performance and power modeling · 0.1FPGA-accelerated architectural simulation · 0.1
YearPublicationVenuePosition
2020 StoneCutter: a very high level instruction set design language
abstract
As the density and capability of reconfigurable computing using FPGAs continues to increase and access to large scale ASIC integration continues to increase, research activities associated with high level synthesis flows have expanded at a similar rate. The goal of these research efforts is to reduce the time and effort required to construct and deploy application-specific architectures. However, these synthesis techniques often force users to consider the entire circuit design space in order to develop a successful implementation. This lack of design specificity often results in hardware design implementations that are difficult to program, difficult to reuse in future designs and make sub-optimal use of hardware resources.
John D. Leidel, David Donofrio, Frank Conlon
CF2
2020 Efficiently Exploiting Low Activity Factors to Accelerate RTL Simulation
abstract
Hardware simulation is a critical tool for design, but its slow speed often bottlenecks the entire design process. Although most signals in a digital design rarely change, most leading simulators still simulate the entirety of the design every cycle. Tracking which signals are unchanged and can thus be reused typically introduces too much overhead to deliver a practical speedup.In this work, we explore the challenge of efficiently detecting opportunities for reuse, and we demonstrate practical techniques to profitably exploit them. Thanks to our novel acyclic partitioning algorithm and other optimizations, our generated simulators outperform open-source and industrial state-of-the-art simulators.
Scott Beamer, David Donofrio
DAC2
2020 DRAM-Less: Hardware Acceleration of Data Processing with New Memory
abstract
General purpose hardware accelerators have become major data processing resources in many computing domains. However, the processing capability of hardware accelerations is often limited by costly software interventions and memory copies to support compulsory data movement between different processors and solid-state drives (SSDs). This in turn also wastes a significant amount of energy in modern accelerated systems. In this work, we propose, DRAM-less, a hardware automation approach that precisely integrates many state-of-the-art phase change memory (PRAM) modules into its data processing network to dramatically reduce unnecessary data copies with a minimum of software modifications. We implement a new memory controller that plugs a real 3x nm multi-partition PRAM to 28nm technology FPGA logic cells and interoperate its design into a real PCIe accelerator emulation platform. The evaluation results reveal that our DRAM-less achieves, on average, 47% better performance than advanced acceleration approaches that use a peer-to-peer DMA.
Jie Zhang 0048, Gyuyoung Park, David Donofrio, John Shalf, Myoungsoo Jung
HPCA3
2020 Evaluating the Numerical Stability of Posit Arithmetic
abstract
The Posit number format has been proposed by John Gustafson as an alternative to the IEEE 754 standard floatingpoint format. Posits offer a unique form of tapered precision whereas IEEE floating-point numbers provide the same relative precision across most of their representational range. Posits are argued to have a variety of advantages including better numerical stability and simpler exception handling.The objective of this paper is to evaluate the numerical stability of Posits for solving linear systems where we evaluate Conjugate Gradient Method to demonstrate an iterative solver and Cholesky-Factorization to demonstrate a direct solver. We show that Posits do not consistently improve stability across a wide range of matrices, but we demonstrate that a simple rescaling of the underlying matrix improves convergence rates for Conjugate Gradient Method and reduces backward error for Cholesky Factorization. We also demonstrate that 16-bit Posit outperforms Float16 for mixed precision iterative refinement - especially when used in conjunction with a recently proposed matrix re-scaling strategy proposed by Nicholas Higham.
Nicholas Buoncristiani, Sanjana Shah, David Donofrio, John Shalf
IPDPS3
2020 Errata to "Exploring Fault-Tolerant Erasure Codes for Scalable All-Flash Array Clusters"
abstract
Presents corrections to author affiliation information in the above mentioned article.
Sungjoon Koh, Jie Zhang 0048, Miryeong Kwon, Jungyeon Yoon, David Donofrio, Nam Sung Kim, Myoungsoo Jung
IEEE Trans. Parallel Distributed Syst.5
2019 TIGER: topology-aware task assignment approach using ising machines
abstract
Optimal mapping of a parallel code's communication graph is increasingly important as both system size and heterogeneity increase. However, the topology-aware task assignment problem is an NP-complete graph isomorphism problem. Existing task scheduling approaches are either heuristic or based on physical optimization algorithms, providing different speed and solution quality tradeoffs. Ising machines such as quantum and digital annealers have recently become available offering an alternative hardware solution to solve certain types of optimization problems. We propose an algorithm that allows expressing the problem for such machines and a domain specific partition strategy that enables to solve larger scale problems. TIGER - topology-aware task assignment mapper tool - implements the proposed algorithm and automatically integrates task - communication graph and an architecture graph into the quantum software environment. We use D-Wave's quantum annealer to demonstrate the solving algorithm and evaluate the proposed tool flow in terms of performance, partition efficiency and solution quality. Results show significant speed-up of the tool flow and reliable solution quality while using TIGER together with the proposed partition.
Anastasiia Butko, George Michelogiannakis, David Donofrio, John Shalf
CF3
2019 Extending classical processors to support future large scale quantum accelerators
abstract
Extensive research in material science together with outstanding engineering efforts allowed quantum technology to be significantly improved hence enabling continuing scaling of quantum circuit size. In around 10 years, quantum annealing circuits have reached 103 qubits and trailing by several years, universal quantum circuits now demonstrate similar trends. From the current trends we can expect that quantum computers will reach thousands of qubits in the next 5--10 years.
Anastasiia Butko, George Michelogiannakis, David Donofrio, John Shalf
CF3
2019 PARADISE - Post-Moore Architecture and Accelerator Design Space Exploration Using Device Level Simulation and Experiments
abstract
An increasing number of technologies are being proposed to preserve digital computing performance scaling as lithographic scaling slows. These technologies include new devices, specialized architectures, memories, and 3D integration. Currently, no end-to-end tool flow is available to rapidly perform architectural-level evaluation using device-level models and for a variety of emerging technologies at once. We propose PARADISE: An open-source comprehensive methodology to evaluate emerging technologies with a vertical simulation flow from the individual device level all the way up to the architec-turallevel. To demonstrate its effectiveness, we use PARADISE to perform end-to-end simulation and analysis of heterogeneous architectures using CNFETs, TFETs, and NCFETs, along with multiple hardware designs. To demonstrate its accuracy, we show that PARADISE has only a 6% mean deviation for delay and 9% for power compared to previous studies using commercial synthesis tools.
Dilip P. Vasudevan, George Michelogiannakis, David Donofrio, John Shalf
ISPASS3
2019 Exploring Fault-Tolerant Erasure Codes for Scalable All-Flash Array Clusters
abstract
Large-scale systems with all-flash arrays have become increasingly common in many computing segments. To make such systems resilient, we can adopt erasure coding such as Reed-Solomon (RS) code as an alternative to replication because erasure coding incurs a significantly lower storage overhead than replication. To understand the impact of using erasure coding on the system performance and other system aspects such as CPU utilization and network traffic, we build a storage cluster that consists of approximately 100 processor cores with more than 50 high-performance solid-state drives (SSDs), and evaluate the cluster with a popular open-source distributed parallel file system, called Ceph. Specifically, we analyze the behaviors of a system adopting erasure coding from the following five viewpoints, and compare with those of another system using replication: (1) storage system I/O performance; (2) computing and software overheads; (3) I/O amplification; (4) network traffic among storage nodes, and (5) impact of physical data layout on performance of RS-coded SSD arrays. For all these analyses, we examine two representative RS configurations, used by Google file systems, and compare them with triple replication employed by a typical parallel file system as a default fault tolerance mechanism. Lastly, we collect 96 block-level traces from the cluster and release them to the public domain for the use of other researchers.
Sungjoon Koh, Jie Zhang 0048, Miryeong Kwon, Jungyeon Yoon, David Donofrio, Nam Sung Kim, Myoungsoo Jung
IEEE Trans. Parallel Distributed Syst.5
2017 OpenSoC system architect: An open toolkit for building soft-cores on FPGAs
abstract
Given the recent difficulty in continuing the classic CMOS manufacturing density and power scaling curves, also known as Moore's Law and Dennard Scaling, respectively, we find that modern complex system architectures are increasingly relying upon accelerators in order to optimize the placement of specific computational workloads. In addition, large-scale computing infrastructures utilized in HPC, data intensive computing, and cloud computing must rely almost exclusively upon commodity device architectures provided by third-party manufacturers. The end result being a final system architecture that lacks specificity for the target software workload. At the same time, there is a trend in the FPGA space of much larger FPGAs with a lot more resources and hardened IP blocks, making this type of architecture design space exploration much easier. The OpenSoC System Architect infrastructure combines several open source design tools and methodologies into a central infrastructure for designing, developing, and verifying the necessary hardware and software modules required to implement application-specific processors for use in FPGAs. The end result is an infrastructure that permits rapid development and deployment of application-specific accelerators and softcores, including a fully functional software development tool chain.
Farzad Fatollahi-Fard, David Donofrio, John Shalf, John D. Leidel, Xi Wang 0009, Yong Chen 0001
FPL2
2017 High-End Computing for Next-Generation Scientific Discovery
Rupak Biswas, David Donofrio, Leonid Oliker
Parallel Comput.2
2016 OpenSoC Fabric: On-chip network generator
abstract
As technology scaling continues, on-chip networks are expected to remain important in future many-core chips due to the increased parallelism and, therefore, communication. However, designing and evaluating large-scale on-chip networks is a nontrivial task given the poor scalability of software simulation for thousands of cores and the intense development effort to develop hardware RTL. In this paper, we describe OpenSoC Fabric. OpenSoC Fabric is a comprehensive on-chip network generator written in Chisel. Chisel generates both C++ and Verilog models from a single code base and has a development effort comparable to functional programming. We describe the internal architecture of OpenSoC Fabric and its powerful list of configuration parameters. We then compare OpenSoC Fabric against pre-validated state-of-the-art simulators using both the generated C++ and Verilog models using FPGAs.
Farzad Fatollahi-Fard, David Donofrio, George Michelogiannakis, John Shalf
ISPASS2
2016 NANDFlashSim: High-Fidelity, Microarchitecture-Aware NAND Flash Memory Simulation
abstract
As the popularity of NAND flash expands in arenas from embedded systems to high-performance computing, a high-fidelity understanding of its specific properties becomes increasingly important. Further, with the increasing trend toward multiple-die, multiple-plane architectures and high-speed interfaces, flash memory systems are expected to continue to scale and cheapen, resulting in their broader proliferation. However, when designing NAND-based devices, making decisions about the optimal system configuration is nontrivial, because flash is sensitive to a number of parameters and suffers from inherent latency variations, and no available tools suffice for studying these nuances. The parameters include the architectures, such as multidie and multiplane, diverse node technologies, bit densities, and cell reliabilities. Therefore, we introduce NANDFlashSim, a high-fidelity, latency-variation-aware, and highly configurable NAND-flash simulator, which implements a detailed timing model for 16 state-of-the-art NAND operations. Using NANDFlashSim, we notably discover the following. First, regardless of the operation, reads fail to leverage internal parallelism. Second, MLC provides lower I/O bus contention than SLC, but contention becomes a serious problem as the number of dies increases. Third, many-die architectures outperform many-plane architectures for disk-friendly workloads. Finally, employing a high-performance I/O bus or an increased page size does not enhance energy savings. Our simulator is available at http://nfs.camelab.org.
Myoungsoo Jung, Wonil Choi, Shuwen Gao, Ellis Herbert Wilson, David Donofrio, John Shalf, Mahmut T. Kandemir
ACM Trans. Storage5
2015 Integrating 3D Resistive Memory Cache into GPGPU for Energy-Efficient Data Processing
abstract
General purpose graphics processing units (GPUs) have become a promising solution to process massive data by taking advantages of multithreading. Thanks to thread-level parallelism, GPU-accelerated applications improve the overall system performance by up to 40 times, compared to CPU-only architecture. However, data-intensive GPU applications often generate large amount of irregular data accesses, which results in cache thrashing and contention problems. The cache thrashing in turn can introduce a large number of off-chip memory accesses, which not only wastes tremendous energy to move data around on-chip cache and off-chip global memory, but also significantly limits system performance due to many stalled load/store instructions. In this work, we redesign the shared last-level cache (LLC) of GPU devices by introducing non-volatile memory (NVM), which can address the cache thrashing issues with low energy consumption. Specifically, we investigate two architectural approaches, one of each employs a 2D planar resistive random-access memory (RRAM) as our baseline NVM-cache and a 3D-stacked RRAM technology. Our baseline NVM-cache replaces the SRAM-based L2 cache with RRAM of similar area size; a memory die consists of eight subarrays, one of which a small fraction of memristor island by constructing 512x512 matrix. Since the feature size of SRAM is around 125 F2 (while that of RRAM around 4 F2), it can offer around 30x bigger storage capacity than the SRAM-based cache. To make our baseline NVM-cache denser, we proposed 3D-stacked NVM-cache, which piles up four memory layers, and each of them has a single pre-decode logic.
Jie Zhang 0048, David Donofrio, John Shalf, Myoungsoo Jung
PACT2
2015 NVMMU: A Non-volatile Memory Management Unit for Heterogeneous GPU-SSD Architectures
abstract
Thanks to massive parallelism in modern Graphics Processing Units (GPUs), emerging data processing applications in GPU computing exhibit ten-fold speedups compared to CPU-only systems. However, this GPU-based acceleration is limited in many cases by the significant data movement overheads and inefficient memory management for host-side storage accesses. To address these shortcomings, this paper proposes a non-volatile memory management unit (NVMMU) that reduces the file data movement overheads by directly connecting the Solid State Disk (SSD) to the GPU. We implemented our proposed NVMMU on a real hardware with commercially available GPU and SSD devices by considering different types of storage interfaces and configurations. In this work, NVMMU unifies two discrete software stacks (one for the SSD and other for the GPU) in two major ways. While a new interface provided by our NVMMU directly forwards file data between the GPU runtime library and the I/O runtime library, it supports non-volatile direct memory access (NDMA) that pairs those GPU and SSD devices via physically shared system memory blocks. This unification in turn can eliminate unnecessary user/kernel-mode switching, improve memory management, and remove data copy overheads. Our evaluation results demonstrate that NVMMU can reduce the overheads of file data movement by 95% on average, improving overall system performance by 78% compared to a conventional IOMMU approach.
Jie Zhang 0048, David Donofrio, John Shalf, Mahmut T. Kandemir, Myoungsoo Jung
PACT2
2015 OpenNVM: An open-sourced FPGA-based NVM controller for low level memory characterization
abstract
Accurate characterization of real device samples is essential for understanding the true potential of the emerging non-volatile memories (NVMs) and identifying their optimal placement in the memory hierarchy. Even though, NVM devices are now available from different manufacturers, lack of an appropriate NVM controller and evaluation platform in the public domain is the main challenge in extracting empirical data from these real devices. In this paper, we present Open-NVM, an open-sourced, highly configurable FPGA based evaluation/characterization platform for various NVM technologies. Through our OpenNVM, this work reveals important low-level NVM characteristics, including i) static and dynamic latency disparity, ii) error rate variation, iii) power consumption behavior, vi) interrelationship between frequency and NVM operational current. In addition, we also examine state-of-the-art write-once-memory (WOM) codes on a real NVM device and study diverse system-level performance impacts based on our findings. All FPGA source code and detailed information of our hardware design is ready to be open-sourced and downloaded for free.
Jie Zhang 0048, Gieseo Park, Mustafa M. Shihab, David Donofrio, John Shalf, Myoungsoo Jung
ICCD4
2012 NANDFlashSim: Intrinsic latency variation aware NAND flash memory system modeling and simulation at microarchitecture level
abstract
As NAND flash memory becomes popular in diverse areas ranging from embedded systems to high performance computing, exposing and understanding flash memory's performance, energy consumption, and reliability becomes increasingly important. Moreover, with an increasing trend towards multiple-die, multiple-plane architectures and high speed interfaces, high performance NAND flash memory systems are expected to continue to scale. This scaling should further reduce costs and thereby widen proliferation of devices based on the technology. However, when designing NAND flash-based devices, making decisions about the optimal system configuration is non-trivial because NAND flash is sensitive to a large number of parameters, and some parameters exhibit significant latency variations. Such parameters include varying architectures such as multi-die and multi-plane, and a host of factors that affect performance, energy consumption, diverse node technology, and reliability. Unfortunately, there are no public domain tools for high-fidelity, microarchitecture level NAND flash memory simulation in existence to assist with making such decisions. Therefore, we introduce NANDFlashSim; a latency variation-aware, detailed, and highly configurable NAND flash simulation model. NANDFlashSim implements a detailed timing model for operations in sixteen state-of-the-art NAND flash operation mode combinations. In addition, NANDFlashSim models energies and reliability of NAND flash memory based on statistics. From our comprehensive experiments using NANDFlashSim, we found that 1) most read cases were unable to leverage the highly-parallel internal architecture of NAND flash regardless of the NAND flash operation mode, 2) the main source of this performance bottleneck is I/O bus activity, not NAND flash activity itself, 3) multi-level-cell NAND flash provides lower I/O bus resource contention than single-level-cell NAND flash, but the resource contention becomes a serious problem as the number of die increases, and 4) preference to employ many dies rather than to employ many planes promises better performance in disk-friendly real workloads. The simulator can be downloaded from http://www.cse.psu.edu/~mqj5086/nfs.
Myoungsoo Jung, Ellis Herbert Wilson, David Donofrio, John Shalf, Mahmut T. Kandemir
MSST3
2011 Hardware/software co-design for energy-efficient seismic modeling
abstract
Reverse Time Migration (RTM) has become the standard for high-quality imaging in the seismic industry. RTM relies on PDE solutions using stencils that are 8th order or larger, which require large-scale HPC clusters to meet the computational demands. However, the rising power consumption of conventional cluster technology has prompted investigation of architectural alternatives that offer higher computational efficiency. In this work, we compare the performance and energy efficiency of three architectural alternatives -- the Intel Nehalem X5530 multicore processor, the NVIDIA Tesla C2050 GPU, and a general-purpose manycore chip design optimized for high-order wave equations called "Green Wave." We have developed an FPGA-accelerated architectural simulation platform to accurately model the power and performance of the Green Wave design. Results show that across highly-tuned high-order RTM stencils, the Green Wave implementation can offer up to 8x and 3.5x energy efficiency improvement per node respectively, compared with the Nehalem and GPU platforms. These results point to the enormous potential energy advantages of our hardware/software co-design methodology.
Jens Krueger, David Donofrio, John Shalf, Marghoob Mohiyuddin, Samuel Williams 0001, Leonid Oliker, Franz-Josef Pfreundt
SC2