Jake Carroll

dblp:209/2781 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-7765-5772ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Institutional Research Computing Capabilities in Australia: 2024
abstract
Institutional research computing infrastructure plays a vital role in Australia’s research ecosystem, complementing and extending national-level facilities. This paper presents an analysis of research computing capabilities across Australian universities and research organisations, examining how institutional infrastructure supports research excellence through localised compute resources, specialised hardware, and cluster solutions. Our study reveals that institutional computing resources of nearly 112,258 CPU cores and 2,241 GPUs serve as essential bridges between desktop computing and national facilities for over 6,000 researchers, enabling research workflows that span from development to large-scale computations. We estimate the total replacement value of this infrastructure to be approximately $144M AUD. Based on detailed infrastructure data provided by research computing facilities across multiple institutions, we identify key patterns in infrastructure deployment, utilisation metrics, and strategic alignment with research priorities. Our findings demonstrate that institutional computing resources not only provide critical support for data-intensive research but also facilitate training and higher-degree research student projects, enable prototyping and development, and ensure data sovereignty compliance when necessary. The analysis shows how these facilities leverage national infrastructure investments while addressing institution-specific needs that cannot be met by national facilities alone. We present evidence that strategic investment in institutional research computing capabilities yields significant returns through increased research productivity, enhanced graduate training, and improved research outcomes. This study provides valuable insights for research organisations planning their computing infrastructure strategies and highlights the importance of maintaining robust institutional computing capabilities alongside national facilities.
Slava Kitaeff, Luc Betbeder-Matibet, Jake Carroll, Stephen Giugni, David Abramson 0001, John Zaitseff, Sarah Walters, David Powell, Chris Bording, Angus Macoustra, Fabien Voisin, Jarrod Hurley
eScience3
2024 An Analysis of Research Data Storage Systems
abstract
The scientific protocols, experiments and instruments that generate data are an integral part of the research lifecycle. Consequently, almost every scientific research institution requires a Research Data Storage System (RDSS). However, RDSS implementations vary significantly due to factors that include cost, geography, workloads, policy, risk tolerance and available technical skills. A RDSS may be on premises, in the public cloud or a mixture of both. Previously we identified 10 key high level features of a RDSS and defined an abstract high-level Research Data Reference Architecture (RDRA). Together, these features enable data fabrics, creating repeatable and consistent structures for low friction, highly efficient data movement and near real time data analysis for decision making and scientific workflows. We build on this earlier work in this paper to present a new Research Data Implementation Architecture (RDIA) that meets the RDRA and can guide implementations without locking in any specific product or service. This paper presents and compares six significant RDSSs and shows how they meet both the RDIA and therefore the RDRA. We identify a new structure – the Research Data Storage System Aggregator (RDSS-A) that describes clusters of RDSSs, and survey five such instances. Finally, we provide a real-world validation of the RDIA by documenting the technical specification of a real RDSS. The new work both clarifies a complex landscape, and aids groups building new systems or adopting existing systems.
Jake Carroll, David Abramson 0001, Bronis R. de Supinski
e-Science1
2023 Why We Need a Reference Architecture for Research Data
abstract
There is little doubt that we have entered an era where data underpins modern science and research in general. In support of this, numerous infrastructures have been designed and built, ranging from proprietary on-prem systems through to distributed commercial clouds. Such implementations provide a range of functions during the research lifecycle from provisioning and cataloguing data assets through to storing and presenting data to computing platforms. In this paper we analyse the underlying principles of such systems and develop a high-level Research Data Reference Architecture (RDRA). Specifically, we identify eight key features of a RDRA that can guide the design, construction, and procurement of implementations without mandating any domain, approach, technical solution, or product choice. As a result, it allows implementers to make local and commercial decisions while still meeting the core requirements of a research data management platform. The intended audience is teams charged with implementing infrastructure in research organizations.
David Abramson 0001, Luc Betbeder-Matibet, Stephen Bird, Jake Carroll, Rhys S. Francis, Wojtek Goscinski, Ai-Lin Soo, Garry Swan, Carmel Walsh, Glenn R. Wightwick, J. Max Wilkinson
e-Science4
2023 Moving small files in a networked environment
Chao Jin 0001, David Abramson 0001, Jake Carroll, Zhengchun Liu, Rajkumar Kettimuthu
Future Gener. Comput. Syst.3
2022 Democratising large scale instrument-based science through e-Infrastructure
abstract
Modern scientific instruments are becoming essential for discoveries because they provide unprecedented insight into physical or biological events – often in real time. However, these instruments may generate large amounts of data, and increasingly they require sophisticated e-infrastructure for analysis, storage and archive. The increasing complexity and scale of the data, processing steps and systems has made it difficult for domain scientists to perform their research, narrowing the user base to a select few. In this paper, we present a framework that democratises large-scale instrument-based science, increasing the number of researchers who can engage. We discuss a prototype at the University of Queensland. The system is illustrated through two case studies, one involving light microscopy imaging of the innate immune system, and the other electron microscopy imaging of the SARS-CoV-2 viral proteins.
David Abramson 0001, Deborah S. Barkauskas, Jake Carroll, Nicholas D. Condon, Naphak Modhiran, James Springfield, Daniel Watterson, Chao Jin 0001
e-Science3
2017 A Metropolitan Area Infrastructure for Data Intensive Science
abstract
The increasing amount of data being collected from simulations, instruments and sensors creates challenges for existing e-Science infrastructure. In particular, it requires new ways of storing, distributing and processing data in order to cope with both the volume and velocity of the data. The University of Queensland has recently designed and deployed MeDiCI, a data fabric that spans the metropolitan area and provides seamless access to data regardless of where it is created, manipulated and archived. MeDiCI is novel in that it exploits temporal and spatial locality to move data on demand in an automated manner. This means that data only needs to reside locally in high speed storage whilst being manipulated, and it can be archived transparently in high capacity, but slower, technologies at other times. MeDiCI is built on commercially available technologies. In this paper, we describe these innovations and present some early results.
David Abramson 0001, Jake Carroll, Chao Jin 0001, Michael Mallon
eScience2