Wayne Kelly

dblp:38/1711 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
1since 2021 · last 2023
0000-0002-8554-4589ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Artificial intelligence and machine learning · 2Software engineering, systems software and programming languages · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Embedded and real-time systems · 39% Processor architecture and microarchitecture · 30% Memory systems · 30%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
many-core architecture
0.312018
CoreVA-MPSoC: A Many-Core Architecture with Tightly Coupled Shared and Local Data Memories · IEEE Trans. Parallel Distributed Syst. 2018
Embedded and real-time systems › embedded hardware platform
MPSoC
0.312018
CoreVA-MPSoC: A Many-Core Architecture with Tightly Coupled Shared and Local Data Memories · IEEE Trans. Parallel Distributed Syst. 2018
Memory systems
shared memory
0.312018
CoreVA-MPSoC: A Many-Core Architecture with Tightly Coupled Shared and Local Data Memories · IEEE Trans. Parallel Distributed Syst. 2018
Embedded and real-time systems
energy-efficient embedded systems
0.112018
CoreVA-MPSoC: A Many-Core Architecture with Tightly Coupled Shared and Local Data Memories · IEEE Trans. Parallel Distributed Syst. 2018

Methods — techniques the papers use, named apart from their topics

post place and route simulation · 0.3
YearPublicationVenuePosition
2023 Examining the efficacy of localised gemcitabine therapy for the treatment of pancreatic cancer using a hybrid agent-based model
abstract
The prognosis for pancreatic ductal adenocarcinoma (PDAC) patients has not significantly improved in the past 3 decades, highlighting the need for more effective treatment approaches. Poor patient outcomes and lack of response to therapy can be attributed, in part, to a lack of uptake of perfusion of systemically administered chemotherapeutic drugs into the tumour. Wet-spun alginate fibres loaded with the chemotherapeutic agent gemcitabine have been developed as a potential tool for overcoming the barriers in delivery of systemically administrated drugs to the PDAC tumour microenvironment by delivering high concentrations of drug to the tumour directly over an extended period. While exciting, the practicality, safety, and effectiveness of these devices in a clinical setting requires further investigation. Furthermore, an in-depth assessment of the drug-release rate from these devices needs to be undertaken to determine whether an optimal release profile exists. Using a hybrid computational model (agent-based model and partial differential equation system), we developed a simulation of pancreatic tumour growth and response to treatment with gemcitabine loaded alginate fibres. The model was calibrated using in vitro and in vivo data and simulated using a finite volume method discretisation. We then used the model to compare different intratumoural implantation protocols and gemcitabine-release rates. In our model, the primary driver of pancreatic tumour growth was the rate of tumour cell division. We were able to demonstrate that intratumoural placement of gemcitabine loaded fibres was more effective than peritumoural placement. Additionally, we quantified the efficacy of different release profiles from the implanted fibres that have not yet been tested experimentally. Altogether, the model developed here is a tool that can be used to investigate other drug delivery devices to improve the arsenal of treatments available for PDAC and other difficult-to-treat cancers in the future.
Adrianne L. Jenner, Wayne Kelly, Michael Dallaston, Robyn Araujo, Isobelle Parfitt, Dominic Steinitz, Pantea Pooladvand, Peter S. Kim, Samantha J. Wade, Kara L. Vine
PLoS Comput. Biol.2
2019 High Resolution Change Detection Using Planet Mosaic
abstract
This paper presents a change detection system using Planet's mosaic dataset. This dataset has higher resolution but fewer bands than data captured from Landsat or Sentinel satellites. Here, an object-based random forest regressor is used to detect vegetation change. The mosaic was separated into individual images, which were then ranked in a list. Results indicate adequate performance with a tradeoff in precision and recall.
Alan Woodley, Connor McLaughlin, Holly Hutson, Shlomo Geva, Timothy Chappell, Wayne Kelly, Dimitri Perrin, Wageeh W. Boles, Lance De Vine
IGARSS6
2018 Scalable Mapping of Streaming Applications onto MPSoCs Using Optimistic Mixed Integer Linear Programming
abstract
Embedded streaming applications are facing increasingly demanding performance requirements in terms of throughput. A common mechanism for providing high compute power with a low energy budget is to use a very large number of low-power cores, often in the form of a Massively Parallel System on Chip (MPSoC). The challenge with programming such massively parallel systems is deciding how to optimally map the computation to individual cores for maximizing throughput. In this work we present an automatic parallelizing compiler for the StreamIt programming language that efficiently and effectively maps computation to individual cores. The compiler must be both effective, meaning that it does a good job of optimizing for throughput; but also efficient, in that the time taken to find such a mapping must scale well as the number of cores and size of the Stream program increases. We improve on previous work that used Integer Linear Programming (ILP) to map StreamIT programs to multicore systems by formulating the mapping problem in a different way using mostly real rather than integer variables. Using so called Mixed Integer Linear Programming (MILP) dramatically reduces the cost compared to standard ILP. This alternative formulation creates what we call an optimistic solution that we then need to adjust slightly to obtain a final feasible solution. We show that this new approach is always close, if not better in terms of effectiveness, while being dramatically better in terms of scalability and efficiency.
Neela Gayen, Johannes Ax, Martin Flasskamp, Christian Klarhorst, Thorsten Jungeblut, Maolin Tang, Wayne Kelly
PDP7
2018 Rapid analysis of metagenomic data using signature-based clustering
abstract
BACKGROUND: Sequencing highly-variable 16S regions is a common and often effective approach to the study of microbial communities, and next-generation sequencing (NGS) technologies provide abundant quantities of data for analysis. However, the speed of existing analysis pipelines may limit our ability to work with these quantities of data. Furthermore, the limited coverage of existing 16S databases may hamper our ability to characterise these communities, particularly in the context of complex or poorly studied environments. RESULTS: In this article we present the SigClust algorithm, a novel clustering method involving the transformation of sequence reads into binary signatures. When compared to other published methods, SigClust yields superior cluster coherence and separation of metagenomic read data, while operating within substantially reduced timeframes. We demonstrate its utility on published Illumina datasets and on a large collection of labelled wound reads sourced from patients in a wound clinic. The temporal analysis is based on tracking the dominant clusters of wound samples over time. The analysis can identify markers of both healing and non-healing wounds in response to treatment. Prominent clusters are found, corresponding to bacterial species known to be associated with unfavourable healing outcomes, including a number of strains of Staphylococcus aureus. CONCLUSIONS: SigClust identifies clusters rapidly and supports an improved understanding of the wound microbiome without reliance on a reference database. The results indicate a promising use for a SigClust-based pipeline in wound analysis and prediction, and a possible novel method for wound management and treatment.
Timothy Chappell, Shlomo Geva, James M. Hogan, Flavia Huygens, Irani U. Rathnayake, Stephen Rudd, Wayne Kelly, Dimitri Perrin
BMC Bioinform.7
2018 CoreVA-MPSoC: A Many-Core Architecture with Tightly Coupled Shared and Local Data Memories
abstract
MPSoCs with hierarchical communication infrastructures are promising architectures for low power embedded systems. Multiple CPU clusters are coupled using an Network-on-Chip (NoC). Our CoreVA-MPSoC targets streaming applications in embedded systems, like signal and video processing. In this work we introduce a tightly coupled shared data memory to each CPU cluster, which can be accessed by all CPUs of a cluster and the NoC with low latency. The main focus is the comparison of different memory architectures and their connection to the NoC. We analyze memory architectures with local data memory only, shared data memory only, and a hybrid architecture integrating both. Implementation results are presented for a 28 nm FD-SOI standard cell technology. A CPU cluster with shared memory shows similar area requirements compared to the local memory architecture. We use post place and route simulations for precise analysis of energy consumption on both cluster and NoC level using the different memory architectures. An architecture with shared data memory shows best performance results in combination with a high resource efficiency. On average, the use of shared memory shows a 17.2 percent higher throughput for a benchmark suite of 10 applications compared to the use of local memory only.
Johannes Ax, Gregor Sievers, Julian Daberkow, Martin Flasskamp, Marten Vohrmann, Thorsten Jungeblut, Wayne Kelly, Mario Porrmann, Ulrich Rückert 0001
IEEE Trans. Parallel Distributed Syst.7
2017 Signature-based clustering for analysis of the wound microbiome
abstract
Chronic wounds present a significant risk to the patient and a substantial drain on health budgets, with the problem likely to worsen markedly with increased incidence of type II diabetes. The wound fluid microbiome is known to influence wound healing outcomes, but is poorly characterised. Next Generation Sequencing approaches yield abundant data from wound samples, but progress in understanding these microbial communities may be hampered by the speed of existing analysis pipelines and limitations on coverage by 16S databases. This paper presents SigClust, a novel clustering method based on binary signatures derived from sequence reads. SigClust yields superior cluster coherence and separation of metagenomic read data in timeframes substantially reduced from those of alternative methods. We demonstrate its utility in the wound context on a preliminary set of labelled patient data. We show how a time course analysis based on tracking the dominant clusters over successive wound samples can identify markers of both successful wound healing and wounds refractory to treatment. Clusters prominent in these analyses are found to correspond to bacterial species known to be implicated as a determinant of wound outcomes, notably a number of strains of Staphylococcus aureus. The clusters obtained rapidly via SigClust support improved understanding of the wound microbiome without direct reliance on a reference database, offering the promise of a SigClust-based pipeline for wound analysis and prediction, and potentially novel methods for wound treatment and management.
Timothy Chappell, Shlomo Geva, James M. Hogan, Flavia Huygens, Wayne Kelly, Dimitri Perrin
BIBM5
2017 A Comparison of Supervised Machine Learning Algorithms for Classification of Communications Network Traffic
Pramitha Perera, Yu-Chu Tian, Colin J. Fidge, Wayne Kelly
ICONIP (1)4
2017 Scalable and efficient data distribution for distributed computing of all-to-all comparison problems
Yi-Fan Zhang 0008, Yu-Chu Tian, Wayne Kelly, Colin J. Fidge
Future Gener. Comput. Syst.3
2016 Data-aware task scheduling for all-to-all comparison problems in heterogeneous distributed systems
Yi-Fan Zhang 0008, Yu-Chu Tian, Colin J. Fidge, Wayne Kelly
J. Parallel Distributed Comput.4
2015 Application of Simulated Annealing to Data Distribution for All-to-All Comparison Problems in Homogeneous Systems
Yi-Fan Zhang 0008, Yu-Chu Tian, Wayne Kelly, Colin J. Fidge, Jing Gao 0006
ICONIP (3)3
2015 Distributed computing of all-to-all comparison problems in heterogeneous systems
abstract
The requirement of distributed computing of all-to-all comparison (ATAC) problems in heterogeneous systems is increasingly important in various domains. Though Hadoop-based solutions are widely used, they are inefficient for the ATAC pattern, which is fundamentally different from the MapReduce pattern for which Hadoop is designed. They exhibit poor data locality and unbalanced allocation of comparison tasks, particularly in heterogeneous systems. The results in massive data movement at runtime and ineffective utilization of computing resources, affecting the overall computing performance significantly. To address these problems, a scalable and efficient data and task distribution strategy is presented in this paper for processing large-scale ATAC problems in heterogeneous systems. It not only saves storage space but also achieves load balancing and good data locality for all comparison tasks. Experiments of bioinformatics examples show that about 89% of the ideal performance capacity of the multiple machines have be achieved through using the approach presented in this paper.
Yi-Fan Zhang 0008, Yu-Chu Tian, Wayne Kelly, Colin J. Fidge
IECON3
2015 Evaluation of interconnect fabrics for an embedded MPSoC in 28 nm FD-SOI
abstract
Embedded many-core architectures contain dozens to hundreds of CPU cores that are connected via a highly scalable NoC interconnect. Our Multiprocessor-System-on-Chip CoreVA-MPSoC combines the advantages of tightly coupled bus-based communication with the scalability of NoC approaches by adding a CPU cluster as an additional level of hierarchy. In this work, we analyze different cluster interconnect implementations with 8 to 32 CPUs and compare them in terms of resource requirements and performance to hierarchical NoCs approaches. Using 28 nm FD-SOI technology the area requirement for 32 CPUs and AXI crossbar is 5.59 mm2including 23.61% for the interconnect at a clock frequency of 830 MHz. In comparison, a hierarchical MPSoC with 4 CPU cluster and 8 CPUs in each cluster requires only 4.83 mm2including 11.61% for the interconnect. To evaluate the performance, we use a compiler for streaming applications to map programs to the different MPSoC configurations. We use this approach for a design-space exploration to find the most efficient architecture and partitioning for an application.
Gregor Sievers, Johannes Ax, Nils Kucza, Martin Flasskamp, Thorsten Jungeblut, Wayne Kelly, Mario Porrmann, Ulrich Rückert 0001
ISCAS6
2014 A distributed computing framework for All-to-All comparison problems
abstract
Distributed computation and storage have been widely used for processing of big data sets. For many big data problems, with the size of data growing rapidly, the distribution of computing tasks and related data can affect the performance of the computing system greatly. In this paper, a distributed computing framework is presented for high performance computing of All-to-All Comparison Problems. A data distribution strategy is embedded in the framework for reduced storage space and balanced computing load. Experiments are conducted to demonstrate the effectiveness of the developed approach. They have shown that about 88% of the ideal performance capacity can be achieved in multiple machines through using the approach presented in this paper.
Yi-Fan Zhang 0008, Yu-Chu Tian, Wayne Kelly, Colin J. Fidge
IECON3
2013 Managing memory and reducing I/O cost for correlation matrix calculation in bioinformatics
abstract
The generation of a correlation matrix from a large set of long gene sequences is a common requirement in many bioinformatics problems such as phylogenetic analysis. The generation is not only computationally intensive but also requires significant memory resources as, typically, few gene sequences can be simultaneously stored in primary memory. The standard practice in such computation is to use frequent input/output (I/O) operations. Therefore, minimizing the number of these operations will yield much faster run-times. This paper develops an approach for the faster and scalable computing of large-size correlation matrices through the full use of available memory and a reduced number of I/O operations. The approach is scalable in the sense that the same algorithms can be executed on different computing platforms with different amounts of memory and can be applied to different problems with different correlation matrix sizes. The significant performance improvement of the approach over the existing approaches is demonstrated through benchmark examples.
Anaththa P. D. Krishnajith, Wayne Kelly, Ross Hayward, Yu-Chu Tian
CIBCB2
2013 Consensus Sigma-70 Promoter Prediction Using Hadoop
abstract
MapReduce frameworks such as Hadoop are well suited to handling large sets of data which can be processed separately and independently, with canonical applications in information retrieval and sales record analysis. Rapid advances in sequencing technology have ensured an explosion in the availability of genomic data, with a consequent rise in the importance of large scale comparative genomics, often involving operations and data relationships which deviate from the classical Map Reduce structure. This work examines the application of Hadoop to patterns of this nature, using as our focus a well established workflow for identifying promoters - binding sites for regulatory proteins - across multiple gene regions and organisms, coupled with the unifying step of assembling these results into a consensus sequence. Our approach demonstrates the utility of Hadoop for problems of this nature, showing how the tyranny of the "dominant decomposition" can be at least partially overcome. It also demonstrates how load balance and the granularity of parallelism can be optimized by pre-processing that splits and reorganizes input files, allowing a wide range of related problems to be brought under the same computational umbrella.
James M. Hogan, Wayne Kelly, Felicity Newell
e-Science2
2010 Using Ownership to Reason about Inherent Parallelism in Object-Oriented Programs
Andrew Craik, Wayne Kelly
CC2
1996 Minimizing Communication While Preserving Parallelism
abstract
All existing methods for automated data/computation decomposition share a common failing: they are very sensitive
Wayne Kelly, William W. Pugh
International Conference on Supercomputing1