VLDB 2026 Research / reviewers in the wild / expert
Matthew Murphy
dblp:65/3433
· DBLP profile ↗
4ranked-venue papers
0as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
GPUs and heterogeneous computing · 77% Cloud and datacenter computing · 23% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › distributed training
large-scale training |
0.5 | 1 | 2021 | Understanding Training Efficiency of Deep Learning Recommendation Models at Scale · HPCA 2021 |
Machine learning › Efficient and distributed learning › distributed training
recommendation model training |
0.5 | 1 | 2021 | Understanding Training Efficiency of Deep Learning Recommendation Models at Scale · HPCA 2021 |
GPUs and heterogeneous computing
GPU-accelerated machine learning |
0.5 | 1 | 2021 | Understanding Training Efficiency of Deep Learning Recommendation Models at Scale · HPCA 2021 |
Methods — techniques the papers use, named apart from their topics
scale-up server design · 1.0embedding table management · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Understanding Training Efficiency of Deep Learning Recommendation Models at ScaleabstractThe use of GPUs has proliferated for machine learning workflows and is now considered mainstream for many deep learning models. Meanwhile, when training state-of-the-art personal recommendation models, which consume the highest number of compute cycles at our large-scale datacenters, the use of GPUs came with various challenges due to having both compute-intensive and memory-intensive components. GPU performance and efficiency of these recommendation models are largely affected by model architecture configurations such as dense and sparse features, MLP dimensions. Furthermore, these models often contain large embedding tables that do not fit into limited GPU memory. The goal of this paper is to explain the intricacies of using GPUs for training recommendation models, factors affecting hardware efficiency at scale, and learnings from a new scale-up GPU server design, Zion. Bilge Acun, Matthew Murphy, Xiaodong Wang 0020, Jade Nie, Carole-Jean Wu, Kim M. Hazelwood |
HPCA | 2 |
| 2006 | Spatiotemporal Denoising and Clustering of fMRI DataabstractThis paper examines combined spatiotemporal denoising and clustering of functional magnetic resonance imaging (fMRI) time series. Most fMRI denoising methods are implemented either in spatial or temporal domain without taking into account both space and time information. In this work, a spatiotemporal denoising method is developed where spatial denoising is implemented by Bayesian shrinkage that uses temporal prior information obtained by statistical testing on all voxel time courses. After the denoising, a set of spatiotemporal features are extracted and characterized by a Gaussian mixture model, which is applied to detect activated areas. The proposed methods have been tested on both synthetic and experimental data, and the results demonstrate their effectiveness. Xiaomu Song, Matthew Murphy, Alice M. Wyrwicz |
ICIP | 2 |
| 2005 | Integrated CORBA Scheduling and Resource Management for Distributed Real-Time Embedded SystemsabstractIntegration of middleware scheduling and resource management services enables open distributed real-time embedded (DRE) applications to meet end-to-end quality of service (QoS) requirements in highly variable operating environments. This paper describes our research on integrating CORBA scheduling and resource management services, and presents experiments we conducted to validate and quantify the benefits of this integration. Our experimental results show that integrating distributed scheduling and resource management in middleware for open DRE systems can offer significant improvements in predictability. Specifically, integrating our stand-alone resource management service with a previously unmanaged experimental baseline application reduced the ratio of missed deadlines from 26% to 10%, and the same application performed even better under the control of integrated scheduling and resource management services, with a missed deadline ratio of only 1%. Kevin Bryan, Lisa Cingiser DiPippo, Victor Fay Wolfe, Matthew Murphy, Jiangyin Zhang, Douglas Niehaus, David Fleeman, David W. Juedes, Lonnie R. Welch, Christopher D. Gill |
IEEE Real-Time and Embedded Technology and Applications Symposium | 4 |
| 2004 | Static Real-Time Data DistributionabstractWe describe the design and implementation of a static real-time data distribution mechanism that uses a real-time event service and a real-time scheduling service to ensure the ontime, and temporally valid delivery of data. The mechanism implements an algorithm, called the just-in-time data distribution algorithm to determine scheduling parameters that will ensure that data that arrives at a requesting target is valid. We use a navy weapons alignment application to demonstrate the necessity and usefulness of the algorithm and implementation. We also present results of tests that demonstrate that our implementation fulfills the guarantees predicted by the theoretical results. Angela Uvarov, Lisa Cingiser DiPippo, Victor Fay Wolfe, Kevin Bryan, Patrick Gadrow, Timothy Henry, Matthew Murphy, Paul R. Work, Louis P. DiPalma |
IEEE Real-Time and Embedded Technology and Applications Symposium | 7 |