Kelly A. Shaw 0001

dblp:84/4440 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5Human-computer interaction and ubiquitous computing · 2Computer networks · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 68% GPUs and heterogeneous computing · 19% Electronic design automation · 10%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache coherence
0.412020
PerpLE: Improving the Speed and Effectiveness of Memory Consistency Testing · MICRO 2020
Memory systems
memory consistency
0.412020
PerpLE: Improving the Speed and Effectiveness of Memory Consistency Testing · MICRO 2020
Electronic design automation
design space exploration
0.212015
GPU Performance and Power Tuning Using Regression Trees · ACM Trans. Archit. Code Optim. 2015
GPUs and heterogeneous computing
GPU performance optimization
0.212015
GPU Performance and Power Tuning Using Regression Trees · ACM Trans. Archit. Code Optim. 2015
Memory systems
cache
0.212014
MRPB: Memory request prioritization for massively parallel processors · HPCA 2014
Memory systems › memory interference
cache contention
0.212014
MRPB: Memory request prioritization for massively parallel processors · HPCA 2014
Memory systems
cache management
0.212014
MRPB: Memory request prioritization for massively parallel processors · HPCA 2014
GPUs and heterogeneous computing › GPU memory
GPU memory hierarchy
0.212014
MRPB: Memory request prioritization for massively parallel processors · HPCA 2014

Methods — techniques the papers use, named apart from their topics

perpetual litmus tests · 0.4litmus testing · 0.4statistical design space exploration · 0.2regression trees · 0.2iterative sampling · 0.2request reordering · 0.2cache bypassing · 0.2
YearPublicationVenuePosition
2020 PerpLE: Improving the Speed and Effectiveness of Memory Consistency Testing
abstract
Even as most of today’s computer systems have turned to parallelism to improve performance, their documentation often remains informal, incomplete or even incorrect regarding their memory consistency models. This leads to programmer and designer confusion and to buggy concurrent systems. Existing tools for empirical memory consistency testing rely on large numbers of iterations of simple multi-threaded litmus tests to perform conformance testing. The current approach typically employs thread synchronization at every iteration, which imposes a significant overhead and can reduce testing performance and efficiency.This paper proposes new litmus test variants called perpetual litmus tests, which allow for consistency testing without periteration synchronization. Perpetual litmus tests use arithmetic sequences in store operations to reduce the required synchronization points. We present PerpLE, a software suite that includes tools for the generation, execution, and analysis of perpetual litmus tests. We introduce an algorithm for determining the outcomes of perpetual litmus tests as well as a scalable linear heuristic algorithm. Evaluating the performance, scalability and ability of our tool to find outcomes of interest on an x86 system, we observe a wider variety of outcomes than litmus7 while experiencing runtime speedups over all litmus7 synchronization modes (8.89x over the default user mode). Compared to the default litmus7 synchronization (user) mode, PerpLE offers over four orders-of-magnitude improvement in the rate with which we detect target outcomes.
Themis Melissaris, Markos Markakis, Kelly A. Shaw 0001, Margaret Martonosi
MICRO3
2020 Analyzing Data Consistency in Smart Home Applications
abstract
Smart Home applications are a quickly growing segment of the Internet of Things (IoT) domain. Many companies offer cloud platforms that enable developers to create and run applications that provide specific Smart Home functions for physical devices installed and connected to the Internet in people's homes. Unfortunately, problems and vulnerabilities relating to data consistency have emerged for these Smart Home platforms, stemming from interfaces that have limited and varied guarantees with respect to atomicity, data consistency, and event ordering. These data consistency concerns are important because many applications are related to devices used for home security and control systems, like smart locks. The Amazon AWS IoT platform enables Internet-connected devices to connect to the AWS Cloud and lets applications in the cloud interact with Internet-connected devices. Because of its many configurable options, the AWS IoT API makes it easy for developers to write applications that do not correctly handle concurrency and data consistency issues, resulting in lost or incorrect state updates. We have developed a static analysis tool for AWS IoT JavaScript applications that detects when software may experience data consistency issues due to how the AWS IoT API was configured and used. We deployed our tool on 10 JavaScript applications in our preliminary analysis; most of these applications could experience erroneous updates of state in the cloud.
Lanqin Wang, Kelly A. Shaw 0001
SIGCSE2
2017 Can we really do it?: Conducting Significant Computer Science Research in Primarily Undergraduate Institutions (PUIs) (Abstract Only)
abstract
Undergraduate research is a critical component of high-quality education in any discipline, including Computer Science (CS). Over the past few years, there has been a dramatic increase in CS undergraduate research activities at colleges and universities, and predominantly undergraduate institutions (PUIs) have an important role to play. Not every university has abundant resources to devote to research, and teaching-focused institutions may face the greatest challenges in this respect. Faculty at PUIs, for example, may face funding and infrastructure challenges and may find themselves stretched thin due to especially high teaching and service expectations. A frequently asked question by new faculty at these institutions is: Is it really possible to conduct meaningful research in such a fast-paced discipline as CS, while juggling a very high teaching and service load? Not only is the answer to this question "Yes!" but there are advantages to conducting research at a non-research institution. Faculty here has access to some of the brightest young minds who will potentially be future graduate students in research-intensive universities. They may have the freedom to do research that is too risky for graduate students. They can work on projects they are interested in, rather than those they know must work. With good time management techniques and careful selection of collaborators and student researchers, faculty here really can conduct important CS research. Thus, the focus of this BOF is to share methods that are helpful in conducting significant and meaningful CS research in a primarily undergraduate or teaching institution.
Farzana Rahman, Suzanne J. Matthews, Kelly A. Shaw 0001, Andrea Pohoreckyj Danyluk
SIGCSE3
2017 Locomotive: Optimizing mobile web traffic using selective compression
abstract
Mobile web traffic and application data demands are growing at a rapid rate and are at odds with resource-constrained, data-capped, wireless mobile devices. Data compression can be used to reduce web traffic, save energy, and make network transfers faster. Compression can, however, hurt performance if not used judiciously. We propose Locomotive, a library that improves the performance of network transfers in wireless mobile networks by employing selective compression based on data type and network conditions. We demonstrate that Locomotive improves performance of web transfers by roughly 12-24% while reducing data usage by 39%.
Themis Melissaris, Kelly A. Shaw 0001, Margaret Martonosi
WoWMoM2
2015 GPU Performance and Power Tuning Using Regression Trees
abstract
GPU performance and power tuning is difficult, requiring extensive user expertise and time-consuming trial and error. To accelerate design tuning, statistical design space exploration methods have been proposed. This article presents Starchart, a novel design space partitioning tool that uses regression trees to approach GPU tuning problems. Improving on prior work, Starchart offers more automation in identifying key design trade-offs and models design subspaces with distinctly different behaviors. Starchart achieves good model accuracy using very few random samples: less than 0.3% of a given design space; iterative sampling can more quickly target subspaces of interest.
Elba Garza, Kelly A. Shaw 0001, Margaret Martonosi
ACM Trans. Archit. Code Optim.3
2014 MRPB: Memory request prioritization for massively parallel processors
abstract
Massively parallel, throughput-oriented systems such as graphics processing units (GPUs) offer high performance for a broad range of programs. They are, however, complex to program, especially because of their intricate memory hierarchies with multiple address spaces. In response, modern GPUs have widely adopted caches, hoping to providing smoother reductions in memory access traffic and latency. Unfortunately, GPU caches often have mixed or unpredictable performance impact due to cache contention that results from the high thread counts in GPUs. We propose the memory request prioritization buffer (MRPB) to ease GPU programming and improve GPU performance. This hardware structure improves caching efficiency of massively parallel workloads by applying two prioritization methods-request reordering and cache bypassing-to memory requests before they access a cache. MRPB then releases requests into the cache in a more cache-friendly order. The result is drastically reduced cache contention and improved use of the limited per-thread cache capacity. For a simulated 16KB L1 cache, MRPB improves the average performance of the entire PolyBench and Rodinia suites by 2.65× and 1.27× respectively, outperforming a state-of-the-art GPU cache management technique.
Kelly A. Shaw 0001, Margaret Martonosi
HPCA2
2013 Starchart: Hardware and software optimization using recursive partitioning regression trees
abstract
Graphics processing units (GPUs) are in increasingly wide use, but significant hurdles lie in selecting the appropriate algorithms, runtime parameter settings, and hardware configurations to achieve power and performance goals with them. Exploring hardware and software choices requires time-consuming simulations or extensive real-system measurements. While some auto-tuning support has been proposed, it is often narrow in scope and heuristic in operation. This paper proposes and evaluates a statistical analysis technique, Starchart, that partitions the GPU hardware/software tuning space by automatically discerning important inflection points in design parameter values. Unlike prior methods, Starchart can identify the best parameter choices within different regions of the space. Our tool is efficient - evaluating at most 0.3% of the tuning space, and often much less - and is robust enough to analyze highly variable real-system measurements, not just simulation. In one case study, we use it to automatically find platform-specific parameter settings that are 6.3× faster (for AMD) and 1.3× faster (for NVIDIA) than a single general setting. We also show how power-optimized parameter settings can save 47W (26% of total GPU power) with little performance loss. Overall, Starchart can serve as a foundation for a range of GPU compiler optimizations, auto-tuners, and programmer tools. Furthermore, because Starchart does not rely on specific GPU features, we expect it to be useful for broader CPU/GPU studies as well.
Kelly A. Shaw 0001, Margaret Martonosi
PACT2
2012 Characterizing and improving the use of demand-fetched caches in GPUs
abstract
Initially introduced as special-purpose accelerators for games and graphics code, graphics processing units (GPUs) have emerged as widely-used high-performance parallel computing platforms. GPUs traditionally provided only software-managed local memories (or scratchpads) instead of demand-fetched caches. Increasingly, however, GPUs are being used in broader application domains where memory access patterns are both harder to analyze and harder to manage in software-controlled caches. In response, GPU vendors have included sizable demand-fetched caches in recent chip designs. Nonetheless, several problems remain. First, since these hardware caches are quite new and highly-configurable, it can be difficult to know when and how to use them; they sometimes degrade performance instead of improving it. Second, since GPU programming is quite distinct from general-purpose programming, application programmers do not yet have solid intuition about which memory reference patterns are amenable to demand-fetched caches.
Kelly A. Shaw 0001, Margaret Martonosi
ICS2
2012 Stargazer: Automated regression-based GPU design space exploration
abstract
Graphics processing units (GPUs) are of increasing interest because they offer massive parallelism for high-throughput computing. While GPUs promise high peak performance, their challenge is a less-familiar programming model with more complex and irregular performance trade-offs than traditional CPUs or CMPs. In particular, modest changes in software or hardware characteristics can lead to large or unpredictable changes in performance. In response to these challenges, our work proposes, evaluates, and offers usage examples of Stargazer1, an automated GPU performance exploration framework based on stepwise regression modeling. Stargazer sparsely and randomly samples parameter values from a full GPU design space and simulates these designs. Then, our automated stepwise algorithm uses these sampled simulations to build a performance estimator that identifies the most significant architectural parameters and their interactions. The result is an application-specific performance model which can accurately predict program runtime for any point in the design space. Because very few initial performance samples are required relative to the extremely large design space, our method can drastically reduce simulation time in GPU studies. For example, we used Stargazer to explore a design space of nearly 1 million possibilities by sampling only 300 designs. For 11 GPU applications, we were able to estimate their runtime with less than 1.1% average error. In addition, we demonstrate several usage scenarios of Stargazer.
Kelly A. Shaw 0001, Margaret Martonosi
ISPASS2