Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Lee W. Howes

dblp:41/2234 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
0since 2021 · last 2015
0009-0009-3491-6387ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 38% GPUs and heterogeneous computing · 24% Parallel and multicore computing · 16%
Software engineering, system software, and programming languages
1 paper
Concurrent programming · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Concurrent programming
memory models
0.212015
HRF-Relaxed: Adapting HRF to the Complexities of Industrial Heterogeneous Memory Models · ACM Trans. Archit. Code Optim. 2015
Memory systems › memory consistency
memory consistency model
0.212015
HRF-Relaxed: Adapting HRF to the Complexities of Industrial Heterogeneous Memory Models · ACM Trans. Archit. Code Optim. 2015
GPUs and heterogeneous computing
GPU computing
0.122010
Performance Comparison of Graphics Processors to Reconfigurable Logic: A Case Study · IEEE Trans. Computers 2010
A comparison of CPUs, GPUs, FPGAs, and massively parallel processor arrays for random number generation · FPGA 2009
Parallel and multicore computing
parallel architecture
0.112009
A comparison of CPUs, GPUs, FPGAs, and massively parallel processor arrays for random number generation · FPGA 2009
Emerging computing paradigms › approximate and stochastic computing › stochastic computing
random number generation
0.112009
A comparison of CPUs, GPUs, FPGAs, and massively parallel processor arrays for random number generation · FPGA 2009

Methods — techniques the papers use, named apart from their topics

scope inclusion · 0.4formalization · 0.4workload characterization · 0.1case study comparison · 0.1monte carlo simulation · 0.1benchmarking · 0.1
YearPublicationVenuePosition
2015 HRF-Relaxed: Adapting HRF to the Complexities of Industrial Heterogeneous Memory Models
abstract
Memory consistency models, or memory models, allow both programmers and program language implementers to reason about concurrent accesses to one or more memory locations. Memory model specifications balance the often conflicting needs for precise semantics, implementation flexibility, and ease of understanding. Toward that end, popular programming languages like Java, C, and C++ have adopted memory models built on the conceptual foundation of Sequential Consistency for Data-Race-Free programs (SC for DRF). These SC for DRF languages were created with general-purpose homogeneous CPU systems in mind, and all assume a single, global memory address space. Such a uniform address space is usually power and performance prohibitive in heterogeneous Systems on Chips (SoCs), and for that reason most heterogeneous languages have adopted split address spaces and operations with nonglobal visibility. There have recently been two attempts to bridge the disconnect between the CPU-centric assumptions of the SC for DRF framework and the realities of heterogeneous SoC architectures. Hower et al. proposed a class of Heterogeneous-Race-Free (HRF) memory models that provide a foundation for understanding many of the issues in heterogeneous memory models. At the same time, the Khronos Group developed the OpenCL 2.0 memory model that builds on the C++ memory model. The OpenCL 2.0 model includes features not addressed by HRF: primarily support for relaxed atomics and a property referred to as scope inclusion. In this article, we generalize HRF to allow formalization of and reasoning about more complicated models using OpenCL 2.0 as a point of reference. With that generalization, we (1) make the OpenCL 2.0 memory model more accessible by introducing a platform for feature comparisons to other models, (2) consider a number of shortcomings in the current OpenCL 2.0 model, and (3) propose changes that could be adopted by future OpenCL 2.0 revisions or by other, related, models.
Benedict R. Gaster, Derek Hower, Lee W. Howes
ACM Trans. Archit. Code Optim.3
2010 Performance Comparison of Graphics Processors to Reconfigurable Logic: A Case Study
abstract
A systematic approach to the comparison of the graphics processor (GPU) and reconfigurable logic is defined in terms of three throughput drivers. The approach is applied to five case study algorithms, characterized by their arithmetic complexity, memory access requirements, and data dependence, and two target devices: the nVidia GeForce 7900 GTX GPU and a Xilinx Virtex-4 field programmable gate array (FPGA). Two orders of magnitude speedup, over a general-purpose processor, is observed for each device for arithmetic intensive algorithms. An FPGA is superior, over a GPU, for algorithms requiring large numbers of regular memory accesses, while the GPU is superior for algorithms with variable data reuse. In the presence of data dependence, the implementation of a customized data path in an FPGA exceeds GPU performance by up to eight times. The trends of the analysis to newer and future technologies are analyzed.
Benjamin Cope, Peter Y. K. Cheung, Wayne Luk, Lee W. Howes
IEEE Trans. Computers4
2009 A comparison of CPUs, GPUs, FPGAs, and massively parallel processor arrays for random number generation
abstract
The future of high-performance computing is likely to rely on the ability to efficiently exploit huge amounts of parallelism. One way of taking advantage of this parallelism is to formulate problems as "embarrassingly parallel" Monte-Carlo simulations, which allow applications to achieve a linear speedup over multiple computational nodes, without requiring a super-linear increase in inter-node communication. However, such applications are reliant on a cheap supply of high quality random numbers, particularly for the three main maximum entropy distributions: uniform, used as a general source of randomness; Gaussian, for discrete-time simulations; and exponential, for discrete-event simulations. In this paper we look at four different types of platform: conventional multi-core CPUs (Intel Core2); GPUs (NVidia GTX 200); FPGAs (Xilinx Virtex-5); and Massively Parallel Processor Arrays (Ambric AM2000). For each platform we determine the most appropriate algorithm for generating each type of number, then calculate the peak generation rate and estimated power efficiency for each device.
David B. Thomas, Lee W. Howes, Wayne Luk
FPGA2
2009 Deriving Efficient Data Movement from Decoupled Access/Execute Specifications
Lee W. Howes, Anton Lokhmotov, Alastair F. Donaldson, Paul H. J. Kelly
HiPEAC1
2006 FPGAs, GPUs and the PS2 - A Single Programming Methodology
abstract
Field programmable gate arrays (FPGAs), graphics processing units (GPUs) and Sony's Playstation 2 vector units offer scope for hardware acceleration of applications. Implementing algorithms on multiple architectures can be a long and complicated process. We demonstrate an approach to compiling for FPGAs, GPUs and PS2 vector units using a unified description based on A Stream Compiler (ASC) for FPGAs. As an example of its use we implement a Monte Carlo simulation using ASC. The unified description allows us to evaluate optimisations for specific architectures on top of a single base description, saving time and effort
Lee W. Howes, Paul Price, Oskar Mencer, Olav Beckmann
FCCM1
2006 Comparing FPGAs to Graphics Accelerators and the Playstation 2 Using a Unified Source Description
abstract
Field programmable gate arrays (FPGAs), graphics processing units (GPUs) and Sony's Playstation 2 vector units offer scope for hardware acceleration of applications. We compare the performance of these architectures using a unified description based onA Stream Compiler(ASC) for FPGAs, which has been extended to target GPUs and PS2 vector units. Programming these architectures from a single description enables us to reason about optimizations for the different architectures. Using the ASC description we implement a Monte Carlo simulation, a fast Fourier transform (FFT) and a weighted sum algorithm. Our results show that without much optimization the GPU is suited to the Monte Carlo simulation, while the weighted sum is better suited to PS2 vector units. FPGA implementations benefit particularly from architecture specific optimizations which ASC allows us to easily implement by adding simple annotations to the shared code.
Lee W. Howes, Paul Price, Oskar Mencer, Olav Beckmann, Oliver Pell
FPL1
2003 Design space exploration with A Stream Compiler
abstract
We consider speeding up general-purpose applications with hardware accelerators. Traditionally hardware accelerators are tediously hand-crafted to achieve top performance ASC (A Stream Complier) simplifies exploration of hardware accelerators by transforming the hardware design task into a software design process using only 'gcc' and 'make' to obtain a hardware netlist. ASC enables programmers to customize hardware accelarators at three levels of abstraction: the architecture level, the functional block level, and the bit level. All three customizations are based on one uniform representation: a single C++ program with custom types and operators for each level of abstraction. This representation allows ASC users to express and reason about the design space, extract parallelism at each level and quickly evaluate different design choices. In addition, since the user has full control over each gate-level resource in the entire design. ASC accelerator performance can always be equal to or better than hand-crafted designs, usually with much less effort. We present several ASC bench marks, including wavelet compression and Kasumi encryption.
Oskar Mencer, David J. Pearce 0001, Lee W. Howes, Wayne Luk
FPT3