Katherine Morrow

dblp:140/2149 · also Katherine (Compton) Morrow · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8Software engineering, systems software and programming languages · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › processing-in-memory
near-DRAM acceleration
0.212015
NDA: Near-DRAM acceleration architecture leveraging commodity DRAM devices and standard memory modules · HPCA 2015
Memory systems › processing-in-memory › memory-centric computing
near-memory accelerator
0.212015
NDA: Near-DRAM acceleration architecture leveraging commodity DRAM devices and standard memory modules · HPCA 2015
Memory systems
processing-in-memory
0.212015
NDA: Near-DRAM acceleration architecture leveraging commodity DRAM devices and standard memory modules · HPCA 2015
Memory systems › DRAM › DRAM architecture
3D-stacked DRAM
0.112015
NDA: Near-DRAM acceleration architecture leveraging commodity DRAM devices and standard memory modules · HPCA 2015
Memory systems
DRAM
0.112015
NDA: Near-DRAM acceleration architecture leveraging commodity DRAM devices and standard memory modules · HPCA 2015

Methods — techniques the papers use, named apart from their topics

through-silicon via · 0.2
YearPublicationVenuePosition
2017 Flexible interconnect in 2.5D ICs to minimize the interposer's metal layers
abstract
In 2.5D ICs, the number of metal layers in the interposer contributes strongly to its manufacturing costs. Many systems implemented as 2.5D ICs include component dies, e.g. FPGA dies, that have flexible interconnect that increase connectivity options within the 2.5D IC. We present the first work to leverage flexible interconnect in FPGA dies within a 2.5D IC to decrease routing metal layers in the interposer. This is done by performing 3D global routing and reassigning flexible pins in the FPGA dies. In our experiments, we reduce the number of metal layers by up to 33% versus the number of layers required before reassigning flexible pins.
Daniel P. Seemuth, Azadeh Davoodi, Katherine Morrow
ASP-DAC3
2015 NDA: Near-DRAM acceleration architecture leveraging commodity DRAM devices and standard memory modules
abstract
Energy consumed for transferring data across the processor memory hierarchy constitutes a large fraction of total system energy consumption, and this fraction has steadily increased with technology scaling. In this paper, we propose near-DRAM acceleration (NDA) architectures, which process data using accelerators 3D-stacked on DRAM devices comprising off-chip main memory modules. NDA transfers most data through high-bandwidth and low-energy 3D interconnects between accelerators and DRAM devices instead of low-bandwidth and high-energy off-chip interconnects between a processor and DRAM devices, substantially reducing energy consumption and improving performance. Unlike previous near-memory processing architectures, NDA is built upon commodity DRAM devices; apart from inserting through-silicon vias (TSVs) to 3D-interconnect DRAM devices and accelerators, NDA requires minimal changes to the commodity DRAM device and standard memory module architectures. This allows NDA to be more easily adopted in both existing and emerging systems. Our experiments demonstrate that, on average, our NDA-based system consumes 46% (68%) lower (data transfer) energy at 1.67× higher performance than a system that integrates the same accelerator logic within the processor itself.
Amin Farmahini Farahani, Jung Ho Ahn, Katherine Morrow, Nam Sung Kim
HPCA3
2015 Guest Editorial FPL 2013
abstract
No abstract available.
João M. P. Cardoso, Pedro C. Diniz, Katherine Morrow
ACM Trans. Reconfigurable Technol. Syst.3
2014 QoS-aware dynamic resource allocation for spatial-multitasking GPUs
abstract
General-purpose computing on GPUs (GPGPU computing) is becoming widely adopted; however, some GPGPU applications fail to fully utilize GPU resources. In these cases, spatial multitasking better exploits the parallelism offered by GPUs by partitioning the GPU resources among simultaneously-running applications. When one or more such applications have quality-of-service (QoS) requirements, enough resources must be allocated for those applications to satisfy their requirements. Remaining resources can be either disabled to reduce power consumption or used to accelerate other applications. However, we observe that the amount of resources for a QoS application to satisfy its performance requirement is dependent in part upon the co-executing applications. In this paper, we propose a runtime technique to dynamically partition GPU resources between concurrently running applications - at least one of which has a QoS requirement. We demonstrate that the proposed technique can satisfy a 100% QoS requirement while also achieving either a 7W power consumption reduction or a 17.57% performance improvement for co-executing best-effort applications.
Paula Aguilera, Katherine Morrow, Nam Sung Kim
ASP-DAC2
2014 Process variation-aware workload partitioning algorithms for GPUs supporting spatial-multitasking
abstract
High-level programming languages have transformed graphics processing units (GPUs) from domain-restricted devices into powerful compute platforms. Yet many “generalpurpose GPU” (GPGPU) applications fail to fully utilize the GPU resources. Executing multiple applications simultaneously on different regions of the GPU (spatial multitasking) thus improves system performance. However, within-die process variations lead to significantly different maximum operating frequencies (Fmax) of the streaming multiprocessors (SMs) within a GPU. As the chip size and number of SMs per chip increase, the frequency variation is also expected to increase, exacerbating the problem. The increased number of SMs also provides a unique opportunity: we can allocate resources to concurrently-executing applications based on how those applications are affected by the different available Fmaxvalues. In this paper, we study the effects of per-SM clocking on spatial multitasking-capable GPUs. We demonstrate two factors that affect the performance of simultaneously-running applications: (i) the SM partitioning algorithm that decides how many resources to assign to each application, and (ii) the assignment of SMs to applications based on the operating frequencies of those SMs and the applications characteristics. Our experimental results show that spatial multitasking that partitions SMs based on application characteristics, when combined with per-SM clocking, can greatly improve application performance by up to 46% on average compared to cooperative multitasking with global clocking.
Paula Aguilera, Jungseob Lee, Amin Farmahini Farahani, Katherine Morrow, Michael J. Schulte, Nam Sung Kim
DATE4
2014 Fair share: Allocation of GPU resources for both performance and fairness
abstract
General-purpose computing on the GPU (GPGPU computing) is becoming widely adopted for an increasing variety of applications. However, it has been shown that as the available computing elements in the GPU increase with every generation some GPGPU applications fail to fully utilize the GPU resources. Spatial multitasking-subdividing GPU resources amongst concurrently-running applications-has been shown to increase overall system performance and utilization for GPGPU computing. However, dividing the computing resources among multiple applications to maximize system performance often results in one application having “unfair” access to GPU resources. Yet, evenly dividing resources among applications does not guarantee equal speedups to each application; nor does it take into account overall system performance. In this paper we examine several different ways to characterize “fairness” for GPGPU spatial multitasking, by balancing individual application's performance and overall system performance. We further present a run-time algorithm to predict and adjust the SM allocation at runtime to meet the desired fairness metric.
Paula Aguilera, Katherine Morrow, Nam Sung Kim
ICCD2
2014 Energy-efficient reconfigurable cache architectures for accelerator-enabled embedded systems
abstract
High-performance embedded systems often include one or more embedded processors tightly coupled with more specialized accelerators. These accelerators improve both performance and energy efficiency because they are specialized for specific (or specific classes of) computations. Data communication between the accelerator and memory, however, is a potential bottleneck for both performance and energy-efficiency. In this paper, we compare and evaluate, for the first time, the impact of L1 data cache design on performance and energy consumption of embedded processor-accelerator systems with shared memory. For this evaluation, we consider data cache design parameters such as size, associativity, and port count, as well as L1 cache sharing between the processor and accelerator. We demonstrate the potential of configurable caches to exploit diversity in cache requirements across hybrid software/hardware applications to significantly improve energy-efficiency while maintaining high performance. Guided by these studies, we propose two techniques for improving energy-efficiency of the cache hierarchy in processor-accelerator systems. The first technique adds configurability to the accelerator-cache interface to allow the accelerator to either share the processor's L1 data cache or use its own private L1 cache. The second technique modifies the L1 cache structure to provide a configurable tradeoff between bandwidth (number of ports) and capacity. Our simulation results show that the first and second techniques improve cache hierarchy energy-efficiency by up to 64% and 33%, respectively, over that of non-configurable caches.
Amin Farmahini Farahani, Nam Sung Kim, Katherine Morrow
ISPASS3
2013 Multi-personality partitioning for heterogeneous systems
abstract
Design flows use graph partitioning both as a precursor to place and route for single devices, and to divide netlists or task graphs among multiple devices. Partitioners have accommodated FPGA heterogeneity via multi-resource constraints, but have not yet exploited the corresponding ability to implement some computations in multiple ways (e.g., LUTs vs. DSP blocks), which could enable a superior solution. This paper introduces multi-personality graph partitioning, which incorporates aspects of resource mapping into partitioning. We present a modified multi-level KLFM partitioning algorithm that also performs heterogeneous resource mapping for nodes with multiple potential implementations (multiple personalities). We evaluate several variants of our multi-personality FPGA circuit partitioner using 21 circuits and benchmark graphs, and show that dynamic resource mapping improves cut size on average by 27% over static mapping for these circuits. We further show that it improves deviation from target resource utilizations by 50% over post-partitioning resource mapping.
Anthony E. Gregerson, Aman Chadha, Katherine Morrow
FPT3
2013 Automated multi-device placement, I/O voltage supply assignment, and pin assignment in circuit board design
abstract
Embedded systems often contain many components, some with multiple Field Programmable Gate Arrays (FPGAs). Designing Printed Circuit Boards (PCBs) for these systems can be a complex process that is often tedious, error-prone, and time-intensive. Existing computer-aided design tools require designers to manually insert components and explicitly define the connections between every component on the PCB - a cumbersome process. A fast PCB design framework requiring reduced designer time and effort would be particularly advantageous for rapid prototyping and short production run PCBs. Therefore, this paper proposes a novel, freely-available open-source framework to capture design intent and automatically implement the design details. Designers express connectivity at a higher level of abstraction than enumerating or drawing each individual trace between components. Given the components and connection requirements, the proposed framework automatically generates component placements, I/O voltage supply assignments, and FPGA pin assignments to minimize trace length. We also propose a novel method to improve trace length estimations during placement, before FPGA pins have actually been assigned to those connections. The proposed framework quickly explores large solution spaces, enabling rapid prototyping and design space exploration, and can lead to lower costs in design time and other non-recurring expenses. We demonstrate that it produces favorable results for various design requirements, which suggests the framework will be especially appreciated by designers of systems with multiple FPGAs having large numbers of flexible pins.
Daniel P. Seemuth, Katherine Morrow
FPT2