Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Matthew A. Watkins

dblp:62/6354 · DBLP profile ↗
← Back
9ranked-venue papers
8as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 6 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Processor architecture and microarchitecture · 52% Reconfigurable computing and FPGAs · 30% Interconnection networks and networks-on-chip · 10%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 50% Software maintenance and evolution · 50%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
dynamic optimization
0.212016
Software transparent dynamic binary translation for coarse-grain reconfigurable architectures · HPCA 2016
Software maintenance and evolution › software evolution › software adaptation
dynamic reconfiguration
0.212016
Software transparent dynamic binary translation for coarse-grain reconfigurable architectures · HPCA 2016
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture
0.212016
Software transparent dynamic binary translation for coarse-grain reconfigurable architectures · HPCA 2016
Processor architecture and microarchitecture › binary translation
dynamic binary translation
0.212016
Software transparent dynamic binary translation for coarse-grain reconfigurable architectures · HPCA 2016
Processor architecture and microarchitecture › multicore design
heterogeneous multicore
0.112010
ReMAP: A Reconfigurable Heterogeneous Multicore Architecture · MICRO 2010
Processor architecture and microarchitecture › chip multiprocessor
inter-core communication
0.112010
ReMAP: A Reconfigurable Heterogeneous Multicore Architecture · MICRO 2010
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable fabric
0.112010
ReMAP: A Reconfigurable Heterogeneous Multicore Architecture · MICRO 2010
Processor architecture and microarchitecture › chip multiprocessor
reconfigurable multicore
0.112010
ReMAP: A Reconfigurable Heterogeneous Multicore Architecture · MICRO 2010
Energy-efficient computing
power-performance tradeoff
0.112016
Software transparent dynamic binary translation for coarse-grain reconfigurable architectures · HPCA 2016
Processor architecture and microarchitecture
chip multiprocessor
0.112006
Leveraging Optical Technology in Future Bus-based Chip Multiprocessors · MICRO 2006
Interconnection networks and networks-on-chip
on-chip interconnect
0.112006
Leveraging Optical Technology in Future Bus-based Chip Multiprocessors · MICRO 2006
Interconnection networks and networks-on-chip
optical interconnection networks
0.112006
Leveraging Optical Technology in Future Bus-based Chip Multiprocessors · MICRO 2006
Memory systems
cache coherence
0.012006
Leveraging Optical Technology in Future Bus-based Chip Multiprocessors · MICRO 2006

Methods — techniques the papers use, named apart from their topics

runtime optimization · 0.5dynamic register information · 0.5
YearPublicationVenuePosition
2018 Characterizing a Commercial Multidimensional Heterogeneous Processor Under GPGPU Workloads
abstract
Single-chip heterogeneous processors are a popular approach to provide power and performance flexibility in today's power constrained world. Popular heterogeneous options include functional CPU+GPU heterogeneity, performance big+little core heterogeneity, and dynamic heterogeneity from DVFS. Commercial systems are now available that incorporate all of these options in a single system. While many past works have looked at one or two dimensions in isolation, very few have looked at all of them in tandem. To best utilize these systems it is important to understand the interaction between workload organization and the different heterogeneous dimensions of a system. This work characterizes the structure of a set of commonly evaluated GPGPU workloads (the Rodinia Benchmarks), evaluates their power and performance on a commercial multidimensional processor, and identifies important behaviors and interactions for each. The results can help understand common workload and processor behavior and guide the design, management, and use of multidimensional heterogeneous systems.
Matthew A. Watkins, Philip Bedoukian
ISPASS1
2017 Characterization of GPGPU workloads on a multidimensional heterogeneous processor
abstract
Systems with multiple forms of heterogeneity, including functional, performance, and dynamic heterogeneity, are now commercially available. The use and tuning of any of these options can impact other options and so it is important to understand their interactions. This work characterizes the power and performance implications of multiple dimensions of heterogeneity from a commercial multidimensional heterogeneous processor on commonly evaluated GPGPU workloads.
Matthew A. Watkins, Philip Bedoukian
ISPASS1
2016 Software transparent dynamic binary translation for coarse-grain reconfigurable architectures
abstract
The end of Dennard Scaling has forced architects to focus on designing for execution efficiency. Course-grained reconfigurable architectures (CGRAs) are a class of architectures that provide a configurable grouping of functional units that aim to bridge the gap between the power and performance of custom hardware and the flexibility of software. Despite their potential benefit, CGRAs face a major adoption challenge as they do not execute a standard instruction stream. Dynamic translation for CGRAs has the potential to solve this problem, but faces non-trivial challenges. Existing attempts either do not achieve the full power and performance potential CGRAs offer or suffer from excessive translation time. In this work we propose DORA, a Dynamic Optimizer for Reconfigurable Architectures, which achieves substantial (2X) power and performance improvements while having low hardware and insertion overhead and benefiting the current execution. In addition to traditional optimizations, DORA leverages dynamic register information to perform optimizations not available to compilers and achieves performance similar to or better than CGRA-targeted compiled code.
Matthew A. Watkins, Tony Nowatzki, Anthony Carno
HPCA1
2010 Dynamically managed multithreaded reconfigurable architectures for chip multiprocessors
abstract
Prior work has demonstrated that reconfigurable logic can significantly benefit certain applications. However, reconfigurable architectures have traditionally suffered from high area overhead and limited application coverage. We present a dynamically managed multithreaded reconfigurable architecture consisting of multiple clusters of shared reconfigurable fabrics that greatly reduces the area overhead of reconfigurability while still offering the same power efficiency and performance benefits. Like other shared SMT and CMP resources, the dynamic partitioning of the reconfigurable resource among sharing threads, along with the co-scheduling of threads among different reconfigurable clusters, must be intelligently managed for the full benefits of the shared fabrics to be realized.
Matthew A. Watkins, David H. Albonesi
PACT1
2010 ReMAP: A Reconfigurable Heterogeneous Multicore Architecture
abstract
This paper presents ReMAP, a reconfigurable architecture geared towards accelerating and parallelizing applications within a heterogeneous CMP. In ReMAP, threads share a common reconfigurable fabric that can be configured for individual thread computation or fine-grained communication with integrated computation. The architecture supports both fine-grained point-to-point communication for pipeline parallelization and fine-grained barrier synchronization. The combination of communication and configurable computation within ReMAP provides the unique ability to perform customized computation while data is transferred between cores, and to execute custom global functions after barrier synchronization. ReMAP demonstrates significantly higher performance and energy efficiency compared to hard-wired communication-only mechanisms, and over what can ideally be achieved by allocating the fabric area to additional or more powerful cores.
Matthew A. Watkins, David H. Albonesi
MICRO1
2009 Revisiting Cache Block Superloading
Matthew A. Watkins, Sally A. McKee, Lambert Schaelicke
HiPEAC1
2008 Shared reconfigurable architectures for CMPS
abstract
This paper investigates reconfigurable architectures suitable for chip multiprocessors (CMPs). Prior research has established that augmenting a conventional processor with reconfigurable logic can dramatically improve the performance of certain application classes, but this comes at non-trivial power and area costs. Given substantial observed time and space differences in fabric usage, we propose that pools of programmable logic should be shared among multiple cores. While a common shared pool is more compact and power efficient, fabric conflicts may lead to large performance losses relative to per-core private fabrics. We identify particular characteristics of past reconfigurable fabric designs that are particularly amenable to fabric sharing. We then propose spatially and temporally shared fabrics in a CMP. The sharing policies that we devise incur negligible performance loss compared to private fabrics, while cutting the area and peak power of the fabric by 4X.
Matthew A. Watkins, Mark J. Cianchetti, David H. Albonesi
FPL1
2007 A Phase-Adaptive Approach to Increasing Cache Performance
Matthew A. Watkins, Sally A. McKee, Lambert Schaelicke
PACT1
2006 Leveraging Optical Technology in Future Bus-based Chip Multiprocessors
abstract
Although silicon optical technology is still in its formative stages, and the more near-term application is chip-to-chip communication, rapid advances have been made in the development of on-chip optical interconnects. In this paper, we investigate the integration of CMOS-compatible optical technology to on-chip cache-coherent buses in future CMPs. While not exhaustive, our investigation yields a hierarchical opto-electrical system that exploits the advantages of optical technology while abiding by projected limitations. Our evaluation shows that, for the applications considered, compared to an aggressive all-electrical bus of similar power and area, significant performance improvements can be achieved using an opto-electrical bus. This performance improvement is largely dependent on the application's bandwidth demand and on the number of implemented wavelengths per optical waveguide. We also present a number of critical areas for future work that we discover in the course of our research
Nevin Kirman, Meyrem Kirman, Rajeev K. Dokania, José F. Martínez, Alyssa B. Apsel, Matthew A. Watkins, David H. Albonesi
MICRO6