Bradley Thwaites

dblp:148/9849 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Emerging computing paradigms · 42% Hardware accelerators and domain-specific architectures · 36% Memory systems · 11%
Artificial intelligence
1 paper
Learning theory · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Emerging computing paradigms
approximate computing
0.942016
RFVP: Rollback-Free Value Prediction with Safe-to-Approximate Loads · ACM Trans. Archit. Code Optim. 2016
Towards Statistical Guarantees in Controlling Quality Tradeoffs for Approximate Acceleration · ISCA 2016
AxGames: Towards Crowdsourcing Quality Target Determination in Approximate Computing · ASPLOS 2016
Hardware accelerators and domain-specific architectures
approximate computing accelerator
0.212016
Towards Statistical Guarantees in Controlling Quality Tradeoffs for Approximate Acceleration · ISCA 2016
Memory systems › memory architecture
memory latency and bandwidth
0.212016
RFVP: Rollback-Free Value Prediction with Safe-to-Approximate Loads · ACM Trans. Archit. Code Optim. 2016
Processor architecture and microarchitecture
value prediction
0.212016
RFVP: Rollback-Free Value Prediction with Safe-to-Approximate Loads · ACM Trans. Archit. Code Optim. 2016
Hardware accelerators and domain-specific architectures
analog computing accelerator
0.212014
General-purpose code acceleration with limited-precision analog computation · ISCA 2014
Hardware accelerators and domain-specific architectures › machine learning accelerator
low-precision arithmetic
0.212014
General-purpose code acceleration with limited-precision analog computation · ISCA 2014
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network acceleration
0.212014
General-purpose code acceleration with limited-precision analog computation · ISCA 2014
Machine learning › Learning theory
statistical guarantees
0.112016
Towards Statistical Guarantees in Controlling Quality Tradeoffs for Approximate Acceleration · ISCA 2016
Compilers and program optimization
code acceleration
0.112014
General-purpose code acceleration with limited-precision analog computation · ISCA 2014

Methods — techniques the papers use, named apart from their topics

statistical optimization · 0.5neural network classifier · 0.5binary classification · 0.5neural network training · 0.4analog circuit design · 0.4rollback-free value prediction · 0.2crowdsourcing · 0.2clopper-pearson exact method · 0.2binomial proportion confidence interval · 0.2approximation · 0.2
YearPublicationVenuePosition
2016 AxGames: Towards Crowdsourcing Quality Target Determination in Approximate Computing
abstract
Approximate computing trades quality of application output for higher efficiency and performance. Approximation is useful only if its impact on application output quality is acceptable to the users. However, there is a lack of systematic solutions and studies that explore users' perspective on the effects of approximation. In this paper, we seek to provide one such solution for the developers to probe and discover the boundary of quality loss that most users will deem acceptable. We propose AxGames, a crowdsourced solution that enables developers to readily infer a statistical common ground from the general public through three entertaining games. The users engage in these games by betting on their opinion about the quality loss of the final output while the AxGames framework collects statistics about their perceptions. The framework then statistically analyzes the results to determine the acceptable levels of quality for a pair of (application, approximation technique). The three games are designed such that they effectively capture quality requirements with various tradeoffs and contexts. To evaluate AxGames, we examine seven diverse applications that produce user perceptible outputs and cover a wide range of domains, including image processing, optical character recognition, speech to text conversion, and audio processing. We recruit 700 participants/users through Amazon's Mechanical Turk to play the games that collect statistics about their perception on different levels of quality. Subsequently, the AxGames framework uses the Clopper-Pearson exact method, which computes a binomial proportion confidence interval, to analyze the collected statistics for each level of quality. Using this analysis, AxGames can statistically project the quality level that satisfies a given percentage of users. The developers can use these statistical projections to tune the level of approximation based on the user experience. We find that the level of acceptable quality loss significantly varies across applications. For instance, to satisfy 90% of users, the level of acceptable quality loss is 2% for one application (image processing) and 26% for another (audio processing). Moreover, the pattern with which the crowd responds to approximation takes significantly different shape and form depending on the class of applications. These results confirm the necessity of solutions that systematically explore the effect of approximation on the end user experience.
Jongse Park, Emmanuel Amaro, Divya Mahajan 0001, Bradley Thwaites, Hadi Esmaeilzadeh
ASPLOS4
2016 Towards Statistical Guarantees in Controlling Quality Tradeoffs for Approximate Acceleration
abstract
Conventionally, an approximate accelerator replaces every invocation of a frequently executed region of code without considering the final quality degradation. However, there is a vast decision space in which each invocation can either be delegated to the accelerator -- improving performance and efficiency -- or run on the precise core -- maintaining quality. In this paper we introduce MITHRA, a co-designed hardware-software solution, that navigates these tradeoffs to deliver high performance and efficiency while lowering the final quality loss. MITHRA seeks to identify whether each individual accelerator invocation will lead to an undesirable quality loss and, if so, directs the processor to run the original precise code. This identification is cast as a binary classification task that requires a cohesive co-design of hardware and software. The hardware component performs the classification at runtime and exposes a knob to the software mechanism to control quality tradeoffs. The software tunes this knob by solving a statistical optimization problem that maximizes benefits from approximation while providing statistical guarantees that final quality level will be met with high confidence. The software uses this knob to tune and train the hardware classifiers. We devise two distinct hardware classifiers, one table-based and one neural network based. To understand the efficacy of these mechanisms, we compare them with an ideal, but infeasible design, the oracle. Results show that, with 95% confidence the table-based design can restrict the final output quality loss to 5% for 90% of unseen input sets while providing 2.5× speedup and 2.6× energy efficiency. The neural design shows similar speedup however, improves the efficiency by 13%. Compared to the table-based design, the oracle improves speedup by 26% and efficiency by 36%. These results show that MITHRA performs within a close range of the oracle and can effectively navigate the quality tradeoffs in approximate acceleration.
Divya Mahajan 0001, Amir Yazdanbakhsh, Jongse Park, Bradley Thwaites, Hadi Esmaeilzadeh
ISCA4
2016 RFVP: Rollback-Free Value Prediction with Safe-to-Approximate Loads
abstract
This article aims to tackle two fundamental memory bottlenecks: limited off-chip bandwidth (bandwidth wall) and long access latency (memory wall). To achieve this goal, our approach exploits the inherent error resilience of a wide range of applications. We introduce an approximation technique, called Rollback-Free Value Prediction (RFVP). When certain safe-to-approximate load operations miss in the cache, RFVP predicts the requested values. However, RFVP does not check for or recover from load-value mispredictions, hence, avoiding the high cost of pipeline flushes and re-executions. RFVP mitigates the memory wall by enabling the execution to continue without stalling for long-latency memory accesses. To mitigate the bandwidth wall, RFVP drops a fraction of load requests that miss in the cache after predicting their values. Dropping requests reduces memory bandwidth contention by removing them from the system. The drop rate is a knob to control the trade-off between performance/energy efficiency and output quality. Our extensive evaluations show that RFVP, when used in GPUs, yields significant performance improvement and energy reduction for a wide range of quality-loss levels. We also evaluate RFVP’s latency benefits for a single core CPU. The results show performance improvement and energy reduction for a wide variety of applications with less than 1% loss in quality.
Amir Yazdanbakhsh, Gennady Pekhimenko, Bradley Thwaites, Hadi Esmaeilzadeh, Onur Mutlu, Todd C. Mowry
ACM Trans. Archit. Code Optim.3
2015 Axilog: language support for approximate hardware design
Amir Yazdanbakhsh, Divya Mahajan 0001, Bradley Thwaites, Jongse Park, Anandhavel Nagendrakumar, Sindhuja Sethuraman, Kartik Ramkrishnan, Nishanthi Ravindran, Rudra Jariwala, Abbas Rahimi, Hadi Esmaeilzadeh, Kia Bazargan
DATE3
2014 Rollback-free value prediction with approximate loads
abstract
This paper demonstrates how to utilize the inherent error resilience of a wide range of applications to mitigate the memory wall -- the discrepancy between core and memory speed. We define a new microarchitecturally-triggered approximation technique called rollback-free value prediction. This technique predicts the value of safe-to-approximate loads when they miss in the cache without tracking mispredictions or requiring costly recovery from misspeculations. This technique mitigates the memory wall by allowing the core to continue computation without stalling for long-latency memory accesses. Our detailed study of the quality trade-offs shows that with a modern out-of-order processor, average 8% (up to 19%) performance improvement is possible with 0.8% (up to 1.8%) average quality loss on an approximable subset of SPEC CPU 2000/2006.
Bradley Thwaites, Gennady Pekhimenko, Hadi Esmaeilzadeh, Amir Yazdanbakhsh, Onur Mutlu, Jongse Park, Girish Mururu, Todd C. Mowry
PACT1
2014 General-purpose code acceleration with limited-precision analog computation
abstract
As improvements in per-transistor speed and energy efficiency diminish, radical departures from conventional approaches are becoming critical to improving the performance and energy efficiency of general-purpose processors. We propose a solution—from circuit to compiler—that enables general-purpose use of limited-precision, analog hardware to accelerate “approximable” code—code that can tolerate imprecise execution. We utilize an algorithmic transformation that automatically converts approximable regions of code from a von Neumann model to an “analog” neural model. We outline the challenges of taking an analog approach, including restricted-range value encoding, limited precision in computation, circuit inaccuracies, noise, and constraints on supported topologies. We address these limitations with a combination of circuit techniques, a hardware/software interface, neural-network training techniques, and compiler support. Analog neural acceleration provides whole application speedup of 3.7× and energy savings of 6.3× with quality loss less than 10% for all except one benchmark. These results show that using limited-precision analog circuits for code acceleration, through a neural approach, is both feasible and beneficial over a range of approximation-tolerant, emerging applications including financial analysis, signal processing, robotics, 3D gaming, compression, and image processing.
Renée St. Amant, Amir Yazdanbakhsh, Jongse Park, Bradley Thwaites, Hadi Esmaeilzadeh, Arjang Hassibi, Luis Ceze, Doug Burger
ISCA4