EDBT 2026 Demo / reviewers in the wild / expert
Kenneth A. Hawick
dblp:01/557 · also Ken A. Hawick
· DBLP profile ↗
33ranked-venue papers
9as first author
0since 2021 · last 2019
0000-0002-2447-3940ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 8 first-authorArtificial intelligence and machine learning · 8 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Parallel and multicore computing · 61% GPUs and heterogeneous computing · 27% High-performance computing · 4% | |
| Computer networks
1 paper |
Internet of things and sensor networks · 100% |
Topics — the 19 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › parallel graph algorithms
connected components |
0.3 | 1 | 2018 | A New Algorithm for Parallel Connected-Component Labelling on GPUs · IEEE Trans. Parallel Distributed Syst. 2018 |
GPUs and heterogeneous computing
GPU computing |
0.3 | 1 | 2018 | A New Algorithm for Parallel Connected-Component Labelling on GPUs · IEEE Trans. Parallel Distributed Syst. 2018 |
Parallel and multicore computing › parallel algorithms › parallel image processing
parallel image analysis |
0.3 | 1 | 2018 | A New Algorithm for Parallel Connected-Component Labelling on GPUs · IEEE Trans. Parallel Distributed Syst. 2018 |
Storage systems
distributed storage |
0.0 | 1 | 2000 | Flexible High-Performance Access to Distributed Storage Resources · HPDC 2000 |
High-performance computing › distributed computing infrastructure
metacomputing |
0.0 | 1 | 1999 | Remote Application Scheduling on Metacomputing Systems · HPDC 1999 |
Distributed systems
distributed data processing |
0.0 | 1 | 1997 | Distributed High Performance Computation for Remote Sensing · SC 1997 |
Parallel and multicore computing
image processing |
0.0 | 1 | 1997 | Distributed High Performance Computation for Remote Sensing · SC 1997 |
Parallel and multicore computing
parallel algorithms |
0.0 | 1 | 1997 | Distributed High Performance Computation for Remote Sensing · SC 1997 |
High-performance computing › iterative methods
conjugate gradient |
0.0 | 1 | 1996 | High-Performance Fortran and Possible Extensions to Support Conjugate Gradient Algorithms · HPDC 1996 |
Parallel and multicore computing › parallel programming models › data-parallel language
high performance fortran |
0.0 | 1 | 1996 | High-Performance Fortran and Possible Extensions to Support Conjugate Gradient Algorithms · HPDC 1996 |
Parallel and multicore computing › parallel computing
parallel programming languages |
0.0 | 1 | 1996 | High-Performance Fortran and Possible Extensions to Support Conjugate Gradient Algorithms · HPDC 1996 |
High-performance computing
scientific computing |
0.0 | 1 | 1996 | High-Performance Fortran and Possible Extensions to Support Conjugate Gradient Algorithms · HPDC 1996 |
Distributed systems › distributed information systems
distributed information management |
0.0 | 1 | 1995 | Distributed Information Management in the National HPCC Software Exchange · SC 1995 |
Storage systems › file systems
distributed file system |
0.0 | 1 | 2000 | Flexible High-Performance Access to Distributed Storage Resources · HPDC 2000 |
Distributed systems
grid computing |
0.0 | 1 | 2000 | Flexible High-Performance Access to Distributed Storage Resources · HPDC 2000 |
Parallel and multicore computing › parallel computing
parallel scientific computing |
0.0 | 1 | 1991 | Scientific modeling with massively parallel SIMD computers · Proc. IEEE 1991 |
Parallel and multicore computing › task allocation
process placement |
0.0 | 1 | 1999 | Remote Application Scheduling on Metacomputing Systems · HPDC 1999 |
Cloud and datacenter computing › cluster resource management and scheduling
resource scheduling |
0.0 | 1 | 1997 | Distributed High Performance Computation for Remote Sensing · SC 1997 |
Parallel and multicore computing › parallel architecture
massively parallel architecture |
0.0 | 1 | 1991 | Scientific modeling with massively parallel SIMD computers · Proc. IEEE 1991 |
Methods — techniques the papers use, named apart from their topics
label-equivalence algorithm · 0.3komura-equivalence algorithm · 0.3distributed object model · 0.0client-server computing · 0.0intrinsic functions · 0.0data distribution directives · 0.0message passing · 0.0recursive algorithm · 0.0process networks · 0.0DRAM future · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Unconventional Exchange: Methods for Statistical Analysis of Virtual GoodsabstractHyperinflation and price volatility in virtual economies has the potential to reduce player satisfaction and decrease developer revenue. This paper describes intuitive analytical methods for monitoring volatility and inflation in virtual economies, with worked examples on the increasingly popular multiplayer game Old School Runescape. Analytical methods drawn from mainstream financial literature are outlined and applied in order to present a high level overview of virtual economic activity of 3467 price series over 180 trading days. Six-monthly volume data for the top 100 most traded items is also used both for monitoring and value estimation, giving a conservative estimate of exchange trading volume of over £60m in real value. Our worked examples show results from a well functioning virtual economy to act as a benchmark for future work. This work contributes to the growing field of virtual economics and game development, describing how data transformations and statistical tests can be used to improve virtual economic design and analysis, with applications in real-time monitoring systems. Oliver James Scholten, Peter I. Cowling, Kenneth A. Hawick, James Alfred Walker |
CoG | 3 |
| 2019 | GraphCombEx: a software tool for exploration of combinatorial optimisation properties of large graphs
David Chalupa, Kenneth A. Hawick |
Soft Comput. | 2 |
| 2018 | A New Algorithm for Parallel Connected-Component Labelling on GPUsabstractConnected-component labelling remains an important and widely-used technique for processing and analysing images and other forms of data in various application areas. Different data sources produce components with different structural features and may be more or less suited to certain connected-component labelling algorithms. Although many efficient serial algorithms exist, determining connected-components on Graphical Processing Units (GPUs) is of interest as many applications use GPUs for processing other parts of the application and labelling on the GPU can avoid expensive memory transfers. The general problem of connected-component labelling is discussed and two existing GPU-based algorithms are discussed-label-equivalence and Komura-equivalence. A new GPU-based, parallel component-labelling algorithm is presented that identifies and eliminates redundant operations in the previous methods for rectilinear two- and three-dimensional datasets. A set of test-cases with a range of structural features and systems sizes is presented and used to evaluate the new labelling algorithm on modern NVIDIA GPU devices and compare it to existing algorithms. The results of the performance evaluation are presented and show that the new algorithm can provide a meaningful performance improvement over previous methods across a range of test cases. Daniel P. Playne, Kenneth A. Hawick |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Computational methods for finding long simple cycles in complex networks
David Chalupa, Phininder Balaghan, Kenneth A. Hawick, Neil A. Gordon |
Knowl. Based Syst. | 3 |
| 2013 | Coarse-to-fine multiclass learning and classification for time-critical domains
Teo Susnjak, Andre L. C. Barczak, Napoleon H. Reyes, Kenneth A. Hawick |
Pattern Recognit. Lett. | 4 |
| 2012 | Comparison of GPU architectures for asynchronous communication with finite-differencing applicationsabstractSUMMARY Graphical processing units (GPUs) are good data‐parallel performance accelerators for solving regular mesh partial differential equations (PDEs) whereby low‐latency communications and high compute to communications ratios can yield very high levels of computational efficiency. Finite‐difference time‐domain methods still play an important role for many PDE applications. Iterative multi‐grid and multilevel algorithms can converge faster than ordinary finite‐difference methods but can be much more difficult to parallelize with GPU memory constraints. We report on some practical algorithmic and data layout approaches and on performance data on a range of GPUs with CUDA. We focus on the use of multiple GPU devices with a single CPU host and the asynchronous CPU/GPU communications issues involved. We obtain more than two orders of magnitude of speedup over a comparable CPU core. Copyright © 2011 John Wiley & Sons, Ltd. Daniel P. Playne, Kenneth A. Hawick |
Concurr. Comput. Pract. Exp. | 2 |
| 2012 | Guest Editor's Introduction: Special Section on Challenges and Solutions in Multicore and Many-Core ComputingabstractIt is our honor to serve as guest editors of this special section of the journal of Concurrency and Computation: Practice and Experience on Frontiers of GPU, Multi- and Many-Core Systems (FGMMS). We are pleased to present nine high-quality contributions in this special issue, where they were first presented at the Frontiers of GPU, Multi- and Many-Core Systems Workshop in conjunction with the 10 th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGrid 2010), held from May 17 to 20, 2010, in Melbourne, Victoria, Australia. The invited papers in this special issue represent augmented works drafted at the beginning of 2010 that addressed the issues below. Multicore and many-core microprocessors are being deployed in a broad spectrum of applications including clusters, clouds, and grids. Both conventional multicore and many-core processors, such as Intel Nehalem and IBM Power7 processors, and unconventional many-core processors, such as NVIDIA Tesla and AMD FireStream graphics processing units (GPUs), hold the promise of increasing performance through parallelism. However, GPU approaches in parallelism are distinctly different from those of conventional multicore and many-core processors, which raises new challenges. For example, how do we optimize applications for conventional multicore and many-core processors? How do we re-engineer applications to take advantage of GPUs’ tremendous computing power in a reasonable cost–benefit ratio? What are effective ways of using GPUs as accelerators? Enormous and rapid progress has been made in accelerator computing over the last two years, but nevertheless we believe the themes developed for the FGMMS workshop and that are represented here in this special issue are still very important ones. In the last two years we have seen the continued rise of GPU computing and indeed its uptake as multi-GPU systems now dominates the top 10 entries within the Top 500 supercomputing systems list 1. Although there has been something of a shakeout of the accelerator technologies that were prevalent three years ago, we have also seen the continued and steady rise of uptake of multicore conventional CPU devices, and an exciting future with combined multicore CPU and GPU devices seems likely. As the articles in this special issue suggest, there are still challenges ahead for application developers to make best use of these future and hybrid highly concurrent systems. There are however still many really important applications that will continue to drive interest, investment, and development of these technologies. The goals of this special issue are to discuss these and other issues and bring together developers of application algorithms and experts in utilizing multicore and many-core processors. We briefly introduce the articles as follows. El Zein and Rendell 2 explore the effect of different GPU programming options (e.g., memory type, memory access methods, and data types) on the performance of routine evaluating sparse matrix vector products and discuss the method for optimal performance. Playne and Hawick 3 report on their approach in accelerating finite-differencing applications using multiple GPU devices with a single CPU host and the asynchronous CPU/GPU communication. Kato and Hosino 4 present their algorithms for speeding up a k-nearest neighbor problem in the recommendation system through multiple GPUs. Ino et al. 5 discuss a cooperative multitasking method for concurrent execution of scientific and graphics applications on GPU and acceleration of compute unified device architecture-based applications using idle GPU cycles in the office. Gillan et al. 6 present a case study on how the instruction-level parallelism offered by three accelerator technologies — field-programmable gate array, GPU, and ClearSpeed — can be exploited in atomic physics with considerable differences in the implementation strategies. Barhen et al. 7 present an unconventional fast Fourier transform implementation scheme for the IBM Cell B.E. processors, named transverse vectorization, and provide the first results for multifast Fourier transform implementation and application on the novel, ultralow power Coherent LogixHyperX processor. Zhou et al. 8 investigate the software and system issues in accelerating climate and weather models in a prototype hybrid computing system, which comprises Intel blades and IBM Cell B.E. blades, connected with both InfiniBand and 1-Gigabit Ethernet and communicate with IBM's Dynamic Application Virtualization software. Qiu and Bae 9 present performance results of two significant bioinformatics applications, gene clustering and dimension reduction, on a Microsoft Windows cluster with up to 768 cores using Message Passing Interface and two variants of threading — Concurrency and Coordination Runtime and Task Parallel Library. Hackenberg et al. 10 present a tool for the graphical program flow analysis of hardware accelerated parallel programs. The tool monitors the hybrid program execution to record and visualize many performance relevant events along the way, and is exemplified through representative real-world applications written for both IBM Cell B.E. processor and NVIDIA Compute Unified Device Architecture (CUDA) API. This work was supported in part by a Microsoft CRMC grant. The guest editors of this special issue would like to express their deep gratitude to all authors, external reviewers, and Geoffrey Fox for their efforts in making this issue possible. Shujia Zhou, Judy Qiu, Kenneth A. Hawick |
Concurr. Comput. Pract. Exp. | 3 |
| 2012 | Adaptive cascade of boosted ensembles for face detection in concept drift
Teo Susnjak, Andre L. C. Barczak, Kenneth A. Hawick |
Neural Comput. Appl. | 3 |
| 2011 | A New Ensemble-Based Cascaded Framework for Multiclass Training with Simple Weak Learners
Teo Susnjak, Andre L. C. Barczak, Napoleon H. Reyes, Kenneth A. Hawick |
CAIP (1) | 4 |
| 2011 | Hypercubic storage layout and transforms in arbitrary dimensions using GPUs and CUDAabstractAbstract Many simulations in the physical sciences are expressed in terms of rectilinear arrays of variables. It is attractive to develop such simulations for use in 1‐, 2‐, 3‐ or arbitrary physical dimensions and also in a manner that supports exploitation of data‐parallelism on fast modern processing devices. We report on data layouts and transformation algorithms that support both conventional and data‐parallel memory layouts. We present our implementations expressed in both conventional serial C code as well as in NVIDIA's Compute Unified Device Architecture concurrent programming language for use on general purpose graphical processing units. We discuss: general memory layouts; specific optimizations possible for dimensions that are powers‐of‐two and common transformations, such as inverting, shifting and crinkling. We present performance data for some illustrative scientific applications of these layouts and transforms using several current GPU devices and discuss the code and speed scalability of this approach. Copyright © 2010 John Wiley & Sons, Ltd. Kenneth A. Hawick, Daniel P. Playne |
Concurr. Comput. Pract. Exp. | 1 |
| 2010 | Asynchronous Communication Schemes for Finite Difference Methods on Multiple GPUsabstractFinite difference methods continue to provide an important and parallelisable approach to many numerical simulations problems. Iterative multigrid and multilevel algorithms can converge faster than ordinary finite difference methods but can be more difficult to parallelise. Data parallel paradigms tend to lend themselves particularly well to solving regular mesh PDEs whereby low latency communications and high compute to communications ratios can yield high levels of computational efficiency and raw performance. We report on some practical algorithmic and data layout approaches and on performance data on a range of Graphical Processing Units (GPUs) with CUDA. We focus on the use of multiple GPU devices with a single CPU host. Daniel P. Playne, Kenneth A. Hawick |
CCGRID | 2 |
| 2010 | Adaptive Ensemble Based Learning in Non-stationary Environments with Variable Concept Drift
Teo Susnjak, Andre L. C. Barczak, Kenneth A. Hawick |
ICONIP (1) | 3 |
| 2010 | Parallel graph component labelling with GPUs and CUDA
Kenneth A. Hawick, Arno Leist, Daniel P. Playne |
Parallel Comput. | 1 |
| 2010 | Spatial pattern growth and emergent animat segregationabstractSpatial agent models can be used to explore self-organizing effects such as pattern growth and segregation. An approximate time line of key animat ideas and agent systems is presented and discussed. These ideas have led to a unique animat simulation Kenneth A. Hawick, Chris Scogings |
Web Intell. Agent Syst. | 1 |
| 2009 | Exploiting graphical processing units for data-parallel scientific applicationsabstractAbstract Graphical processing units (GPUs) have recently attracted attention for scientific applications such as particle simulations. This is partially driven by low commodity pricing of GPUs but also by recent toolkit and library developments that make them more accessible to scientific programmers. We discuss the application of GPU programming to two significantly different paradigms—regular mesh field equations with unusual boundary conditions and graph analysis algorithms. The differing optimization techniques required for these two paradigms cover many of the challenges faced when developing GPU applications. We discuss the relevance of these application paradigms to simulation engines and games. GPUs were aimed primarily at the accelerated graphics market but since this is often closely coupled to advanced game products it is interesting to speculate about the future of fully integrated accelerator hardware for both visualization and simulation combined. As well as reporting the speed‐up performance on selected simulation paradigms, we discuss suitable data‐parallel algorithms and present code examples for exploiting GPU features like large numbers of threads and localized texture memory. We find a surprising variation in the performance that can be achieved on GPUs for our applications and discuss how these findings relate to past known effects in parallel computing such as memory speed‐related super‐linear speed up. Copyright © 2009 John Wiley & Sons, Ltd. Arno Leist, Daniel P. Playne, Kenneth A. Hawick |
Concurr. Comput. Pract. Exp. | 3 |
| 2008 | Altruism Amongst Spatial Predator-Prey Animats
Chris Scogings, Kenneth A. Hawick |
ALIFE | 2 |
| 2008 | Gaining colour stability in live image capturingabstractDigital colour cameras are dramatically falling in price, making them affordable for ubiquitous appliances in many applications. An attempt to use colour information reveals a significant problem that usually escapes our awareness. Due to the adaptive nature of the human visual system in most cases we do not recognise most changes in illumination characteristics, a camera however will measure scenes under changing illumination differently. Attempts to deduce object colour from the images will need to cope with the influence of the illumination and the camera's characteristics. Furthermore, a large variety of colour spaces are available to describe colour. Differences between them and their fitness to quantify colour are discussed. This paper tries to establish a basic understanding of the intricacies behind the processes involved in capturing images and recognising colour- from light as a stimulus to the sensed colour values in cameras. The goal is to outline a novel approach fusing common industrial best practices with dynamic adaptation capabilities needed for robustly measuring colour using cameras in real-time. First positive results towards improving colour based reasoning on adaptable colour spaces are stated as an outlook for further development directions. Guy K. Kloss, Napoleon H. Reyes, Martin J. Johnson, Kenneth A. Hawick |
ICARCV | 4 |
| 2005 | Scientific Data Management in a Grid Environment
Heath A. James, Kenneth A. Hawick |
J. Grid Comput. | 2 |
| 2003 | Distributed frameworks and parallel algorithms for processing large-scale geographic data
Kenneth A. Hawick, Paul D. Coddington, Heath A. James |
Parallel Comput. | 1 |
| 2001 | Dynamic Cluster Configuration and Management using JavaSpacesabstractIntroduction Like many organisations we are interested in managing computer clusters as resources for carrying out computationally intensive tasks. In common with many University organisations our resources are heterogeneous and have to serve many purposes. Although we have been able to develop computational clusters where a tranche of nodes was entirely dedicated we believe it is highly attractive to support cluster management of resources that may fluctuate in number or availability at different times. This is important for management of local resources but is also an important foundation idea for sharing resources between different organisations. In this paper we report on our prototype cluster management framework using a tuple spaces model and show how a coordination management system can be set up independent of any communications mechanism needed for achieving high performance amongst nodes working on parallel jobs. The tuple space model widely popularised by Gelernter Kenneth A. Hawick, Heath A. James |
CLUSTER | 1 |
| 2001 | Topic 16: Cluster Computing
Mark Baker, John M. Brooke, Kenneth A. Hawick, Rajkumar Buyya |
Euro-Par | 3 |
| 2001 | Topic 19: Problem Solving Environments
David W. Walker, Kenneth A. Hawick, Domenico Laforenza, Efstratios Gallopoulos |
Euro-Par | 2 |
| 2001 | Resource Discovery for Dynamic Clusters in Computational GridsabstractWe describe a de-centralised approach to resource management and discovery, based on a community of interacting software agents. Each agent either represents a user application, a resource, or a MatchMaking service. The proposed approach can support dynamic registration of resources and user tasks, facilitating the establishment of dynamic clusters. Resource capability and task requirements are described using an object based data model, enabling new types of devices or new features in existing devices to be identified. A comparison with the Discovery and LookUp services in Jini and TSpaces is also provided. Omer F. Rana, Daniel Bunford-Jones, Kenneth A. Hawick, David W. Walker, Matthew Addis, Mike Surridge |
IPDPS | 3 |
| 2001 | Asynchronous Transfer Mode and other Network Technologies for Wide-Area and High-Performance Cluster Computing
Kenneth A. Hawick, Heath A. James |
J. Supercomput. | 1 |
| 2000 | Agent based Resource Discovery for Dynamic Clusters
Omer F. Rana, Daniel Bunford-Jones, David W. Walker, Matthew Addis, Mike Surridge, Kenneth A. Hawick |
CLUSTER | 6 |
| 2000 | Flexible High-Performance Access to Distributed Storage ResourcesabstractDescribes a software architecture for storage services in computational grid environments. Based upon a lightweight message-passing paradigm, the architecture enables the provision and composition of active, distributed storage services. These services can then cooperatively provide access to distributed storage in a manner potentially optimized for dataset and resource environments. We report on the design and implementation of a distributed file system and a dataset-specific satellite imagery service using the architecture. We discuss data movement and storage issues and implications for future work with the architecture. Craig J. Patten, Kenneth A. Hawick |
HPDC | 2 |
| 1999 | Remote Application Scheduling on Metacomputing SystemsabstractEfficient and robust metacomputing requires the decomposition of complex jobs into tasks that must be scheduled on distributed processing nodes. There are various ways of creating a schedule and implementing it efficiently, depending upon global system state knowledge. Many computations may be structured as process networks, where data is either pushed from the source node to the target node where it will be used, or is pulled from source to target at the instigation of the target. We have developed a metacomputing infrastructure to investigate this idea, which employs the concept of a rich data pointer, the DISCWorld Remote Access Mechanism (DRAM), which can point to either data or services, and can be traded in a client/multiple-server model. We present an extension of the DRAM concept and implementation to represent and describe data that has not yet been created, the "DRAM Future" (DRAMF). We show how the use of the DRAMF facilitates efficient metacomputing scheduling and runtime optimisation on high performance distributed systems. We present a recursive algorithm for determining the optimal placement of a job's components in the presence of partial system state information. This algorithm uses only a selected subset of all available processing nodes, and we implement it using DRAMFs. There are many research issues to consider when designing a robust and general algorithm for scheduling and process placement on distributed systems. We address some of these issues in our Distributed Information Systems Control World project[6] as do other research projects[2]. Heath A. James, Kenneth A. Hawick |
HPDC | 2 |
| 1999 | Interfacing to distributed active data archives
Kenneth A. Hawick, Paul D. Coddington |
Future Gener. Comput. Syst. | 1 |
| 1999 | DISCWorld: an environment for service-based matacomputing
Kenneth A. Hawick, Heath A. James, A. J. Silis, Duncan A. Grove, Craig J. Patten, J. A. Mathew, Paul D. Coddington, K. E. Kerry, J. F. Hercus, F. A. Vaughan |
Future Gener. Comput. Syst. | 1 |
| 1997 | Distributed High Performance Computation for Remote SensingabstractWe describe distributed and parallel algorithms for processing remotely sensed data such as geostationary satellite imagery. We have built a distributed data repository based around the client-server computing model across wide-area ATM networks, with embedded parallel and high-performance processing modules. We focus on algorithms for classification, geo-rectification, correlation and histogram analysis of the data. We consider characteristics of image data collected from the Japanese GMS5 geostationary meteorological satellite, and some analysis techniques we have applied to it. As well as providing a browsing interface to our data collection, our system provides processing and analysis services on-demand. We are developing our system to carry out processing and data reduction services at-a-distance, enabling remote users with limited bandwidth, access to our system, to obtain useful derived data products at the resolution they require.Our target hardware consists of a heterogeneous collection of distributed workstations, multi-processor servers and massively parallel computers at locations throughout Australia. These platforms are connected by ATM-based LANs and also through ATM switches across long distance WANs such as Telstra's Experimental Broadband Network, connecting Adelaide, Melbourne and Canberra. Our particular interest in constructing remote data access and processing services is the potential to utilise such resources as are available to a given user, yielding the best performance compromise of data processing and data delivery. To this end, we are building a set of resource scheduling and management utilities that will integrate the processing modules we describe. We have considered a number of software frameworks for building our integrated system and are focusing on a distributed object model using WWW protocols and the Java language. Kenneth A. Hawick, Heath A. James |
SC | 1 |
| 1996 | High-Performance Fortran and Possible Extensions to Support Conjugate Gradient AlgorithmsabstractEvaluates the High Performance Fortran (HPF) language for the compact expression and efficient implementation of conjugate-gradient iterative matrix-solvers on high-performance computing and communications (HPCC) platforms. We discuss the use of intrinsic functions, data distribution directives and explicitly parallel constructs to optimize performance by minimizing communications requirements in a portable manner. We focus on implementations using the existing HPF definitions but also discuss issues arising that may influence a revised definition for HPF-2. Some of the codes discussed are available on the World Wide Web at http://www.npac.syr.edu/hpfa/, along with other educational and discussion material related to applications in HPF. Kivanç Dinçer, Geoffrey C. Fox, Kenneth A. Hawick |
HPDC | 3 |
| 1995 | Distributed Information Management in the National HPCC Software Exchange
Shirley Browne, Jack J. Dongarra, Geoffrey C. Fox, Kenneth A. Hawick, Ken Kennedy, Rick L. Stevens, Robert Olson, Tom Rowan |
SC | 4 |
| 1991 | Scientific modeling with massively parallel SIMD computersabstractA number of scientific models are discussed that possess a high degree of inherent parallelism. For simulation purposes this is exploited by employing a massively parallel SIMD (single instruction multiple data) computer. The authors describe one such computer, the distributed array processor (DAP), and discuss the optimal mapping of a typical problem onto the computer architecture to best exploit the model parallelism. By focusing on specific models currently under study, they exemplify the types of problems which benefit most from a parallel implementation. The extent of this benefit is considered relative to implementation on a machine of conventional architecture.> Nigel B. Wilding, Arthur S. Trew, Kenneth A. Hawick, G. Stuart Pawley |
Proc. IEEE | 3 |