VLDB 2026 Research / reviewers in the wild / expert
Matthew A. Watkins
dblp:62/6354
· DBLP profile ↗
9ranked-venue papers
8as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 6 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Processor architecture and microarchitecture · 52% Reconfigurable computing and FPGAs · 30% Interconnection networks and networks-on-chip · 10% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 50% Software maintenance and evolution · 50% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
dynamic optimization |
0.2 | 1 | 2016 | Software transparent dynamic binary translation for coarse-grain reconfigurable architectures · HPCA 2016 |
Software maintenance and evolution › software evolution › software adaptation
dynamic reconfiguration |
0.2 | 1 | 2016 | Software transparent dynamic binary translation for coarse-grain reconfigurable architectures · HPCA 2016 |
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture |
0.2 | 1 | 2016 | Software transparent dynamic binary translation for coarse-grain reconfigurable architectures · HPCA 2016 |
Processor architecture and microarchitecture › binary translation
dynamic binary translation |
0.2 | 1 | 2016 | Software transparent dynamic binary translation for coarse-grain reconfigurable architectures · HPCA 2016 |
Processor architecture and microarchitecture › multicore design
heterogeneous multicore |
0.1 | 1 | 2010 | ReMAP: A Reconfigurable Heterogeneous Multicore Architecture · MICRO 2010 |
Processor architecture and microarchitecture › chip multiprocessor
inter-core communication |
0.1 | 1 | 2010 | ReMAP: A Reconfigurable Heterogeneous Multicore Architecture · MICRO 2010 |
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable fabric |
0.1 | 1 | 2010 | ReMAP: A Reconfigurable Heterogeneous Multicore Architecture · MICRO 2010 |
Processor architecture and microarchitecture › chip multiprocessor
reconfigurable multicore |
0.1 | 1 | 2010 | ReMAP: A Reconfigurable Heterogeneous Multicore Architecture · MICRO 2010 |
Energy-efficient computing
power-performance tradeoff |
0.1 | 1 | 2016 | Software transparent dynamic binary translation for coarse-grain reconfigurable architectures · HPCA 2016 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 1 | 2006 | Leveraging Optical Technology in Future Bus-based Chip Multiprocessors · MICRO 2006 |
Interconnection networks and networks-on-chip
on-chip interconnect |
0.1 | 1 | 2006 | Leveraging Optical Technology in Future Bus-based Chip Multiprocessors · MICRO 2006 |
Interconnection networks and networks-on-chip
optical interconnection networks |
0.1 | 1 | 2006 | Leveraging Optical Technology in Future Bus-based Chip Multiprocessors · MICRO 2006 |
Memory systems
cache coherence |
0.0 | 1 | 2006 | Leveraging Optical Technology in Future Bus-based Chip Multiprocessors · MICRO 2006 |
Methods — techniques the papers use, named apart from their topics
runtime optimization · 0.5dynamic register information · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Characterizing a Commercial Multidimensional Heterogeneous Processor Under GPGPU WorkloadsabstractSingle-chip heterogeneous processors are a popular approach to provide power and performance flexibility in today's power constrained world. Popular heterogeneous options include functional CPU+GPU heterogeneity, performance big+little core heterogeneity, and dynamic heterogeneity from DVFS. Commercial systems are now available that incorporate all of these options in a single system. While many past works have looked at one or two dimensions in isolation, very few have looked at all of them in tandem. To best utilize these systems it is important to understand the interaction between workload organization and the different heterogeneous dimensions of a system. This work characterizes the structure of a set of commonly evaluated GPGPU workloads (the Rodinia Benchmarks), evaluates their power and performance on a commercial multidimensional processor, and identifies important behaviors and interactions for each. The results can help understand common workload and processor behavior and guide the design, management, and use of multidimensional heterogeneous systems. Matthew A. Watkins, Philip Bedoukian |
ISPASS | 1 |
| 2017 | Characterization of GPGPU workloads on a multidimensional heterogeneous processorabstractSystems with multiple forms of heterogeneity, including functional, performance, and dynamic heterogeneity, are now commercially available. The use and tuning of any of these options can impact other options and so it is important to understand their interactions. This work characterizes the power and performance implications of multiple dimensions of heterogeneity from a commercial multidimensional heterogeneous processor on commonly evaluated GPGPU workloads. Matthew A. Watkins, Philip Bedoukian |
ISPASS | 1 |
| 2016 | Software transparent dynamic binary translation for coarse-grain reconfigurable architecturesabstractThe end of Dennard Scaling has forced architects to focus on designing for execution efficiency. Course-grained reconfigurable architectures (CGRAs) are a class of architectures that provide a configurable grouping of functional units that aim to bridge the gap between the power and performance of custom hardware and the flexibility of software. Despite their potential benefit, CGRAs face a major adoption challenge as they do not execute a standard instruction stream. Dynamic translation for CGRAs has the potential to solve this problem, but faces non-trivial challenges. Existing attempts either do not achieve the full power and performance potential CGRAs offer or suffer from excessive translation time. In this work we propose DORA, a Dynamic Optimizer for Reconfigurable Architectures, which achieves substantial (2X) power and performance improvements while having low hardware and insertion overhead and benefiting the current execution. In addition to traditional optimizations, DORA leverages dynamic register information to perform optimizations not available to compilers and achieves performance similar to or better than CGRA-targeted compiled code. Matthew A. Watkins, Tony Nowatzki, Anthony Carno |
HPCA | 1 |
| 2010 | Dynamically managed multithreaded reconfigurable architectures for chip multiprocessorsabstractPrior work has demonstrated that reconfigurable logic can significantly benefit certain applications. However, reconfigurable architectures have traditionally suffered from high area overhead and limited application coverage. We present a dynamically managed multithreaded reconfigurable architecture consisting of multiple clusters of shared reconfigurable fabrics that greatly reduces the area overhead of reconfigurability while still offering the same power efficiency and performance benefits. Like other shared SMT and CMP resources, the dynamic partitioning of the reconfigurable resource among sharing threads, along with the co-scheduling of threads among different reconfigurable clusters, must be intelligently managed for the full benefits of the shared fabrics to be realized. Matthew A. Watkins, David H. Albonesi |
PACT | 1 |
| 2010 | ReMAP: A Reconfigurable Heterogeneous Multicore ArchitectureabstractThis paper presents ReMAP, a reconfigurable architecture geared towards accelerating and parallelizing applications within a heterogeneous CMP. In ReMAP, threads share a common reconfigurable fabric that can be configured for individual thread computation or fine-grained communication with integrated computation. The architecture supports both fine-grained point-to-point communication for pipeline parallelization and fine-grained barrier synchronization. The combination of communication and configurable computation within ReMAP provides the unique ability to perform customized computation while data is transferred between cores, and to execute custom global functions after barrier synchronization. ReMAP demonstrates significantly higher performance and energy efficiency compared to hard-wired communication-only mechanisms, and over what can ideally be achieved by allocating the fabric area to additional or more powerful cores. Matthew A. Watkins, David H. Albonesi |
MICRO | 1 |
| 2009 | Revisiting Cache Block Superloading
Matthew A. Watkins, Sally A. McKee, Lambert Schaelicke |
HiPEAC | 1 |
| 2008 | Shared reconfigurable architectures for CMPSabstractThis paper investigates reconfigurable architectures suitable for chip multiprocessors (CMPs). Prior research has established that augmenting a conventional processor with reconfigurable logic can dramatically improve the performance of certain application classes, but this comes at non-trivial power and area costs. Given substantial observed time and space differences in fabric usage, we propose that pools of programmable logic should be shared among multiple cores. While a common shared pool is more compact and power efficient, fabric conflicts may lead to large performance losses relative to per-core private fabrics. We identify particular characteristics of past reconfigurable fabric designs that are particularly amenable to fabric sharing. We then propose spatially and temporally shared fabrics in a CMP. The sharing policies that we devise incur negligible performance loss compared to private fabrics, while cutting the area and peak power of the fabric by 4X. Matthew A. Watkins, Mark J. Cianchetti, David H. Albonesi |
FPL | 1 |
| 2007 | A Phase-Adaptive Approach to Increasing Cache Performance
Matthew A. Watkins, Sally A. McKee, Lambert Schaelicke |
PACT | 1 |
| 2006 | Leveraging Optical Technology in Future Bus-based Chip MultiprocessorsabstractAlthough silicon optical technology is still in its formative stages, and the more near-term application is chip-to-chip communication, rapid advances have been made in the development of on-chip optical interconnects. In this paper, we investigate the integration of CMOS-compatible optical technology to on-chip cache-coherent buses in future CMPs. While not exhaustive, our investigation yields a hierarchical opto-electrical system that exploits the advantages of optical technology while abiding by projected limitations. Our evaluation shows that, for the applications considered, compared to an aggressive all-electrical bus of similar power and area, significant performance improvements can be achieved using an opto-electrical bus. This performance improvement is largely dependent on the application's bandwidth demand and on the number of implemented wavelengths per optical waveguide. We also present a number of critical areas for future work that we discover in the course of our research Nevin Kirman, Meyrem Kirman, Rajeev K. Dokania, José F. Martínez, Alyssa B. Apsel, Matthew A. Watkins, David H. Albonesi |
MICRO | 6 |