VLDB 2026 Research / reviewers in the wild / expert
Jake Tronge
dblp:311/8904 · also Jacob Tronge
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2024
0000-0002-4008-8719ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | An HPC-Container Based Continuous Integration Tool for Detecting Scaling and Performance Issues in HPC ApplicationsabstractTesting is one of the most important steps in software development–it ensures the quality of software. Continuous Integration (CI) is a widely used testing standard that can report software quality to the developer in a timely manner during development progress. Performance, especially scalability, is another key factor for High Performance Computing (HPC) applications. There are many existing profiling and performance tools for HPC applications, but none of these are integrated into CI tools. In this work, we propose BeeSwarm, an HPC container based parallel scaling performance system that can be easily applied to the current CI test environments. BeeSwarm is mainly designed for HPC application developers who need to monitor how their applications can scale on different compute resources. We demonstrate BeeSwarm using three different HPC applications: CoMD, LULESH and NWChem. We utilize GitHub Actions and provision resources from Google Compute Engine. Our results show that BeeSwarm can be used for scalability and performance testing of a variety of HPC applications, allowing developers to monitor application performance over time. Jake Tronge, Jieyang Chen, Patricia Grubel, Tim Randles, Rusty Davis, Quincy Wofford, Steven Anaya, Qiang Guan |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | Improving MPI Safety for Modern LanguagesabstractA program or library is considered safe when it’s guaranteed that programmer error cannot cause undefined behavior. MPI, both the standard and its implementations, is not designed to be type or memory safe. The standard requires that types must match across point-to-point communications and that collective call arguments must be the same across all ranks; it is up to the user in most cases to avoid these errors, and if not caught, then program behavior is undefined. Existing research attempts to help application developers find these errors using profiling and debugging tools, but these do not focus on solving the problems of safety. These tools are usually designed to catch errors during development but safety errors can depend on the environment, hardware, and input values, thus they are likely to show up even during production runs. To properly bind and use MPI, many modern memory safe languages, such as Rust, have to carefully validate and check for these errors, making it hard to achieve the performance of languages like Fortran and C. Our work examines how to improve MPI safety, both type and memory safety, at the implementation level and within programming languages; in this way it is possible to ensure valid communication and safety whether in development or in production. Our work presents safe point-to-point messaging prototypes within an MPI implementation and as a UCX-based library written in the Rust programming language. These prototypes are both designed to catch point-to-point communication errors at runtime, specifically datatype mismatches. We analyze results from these prototypes to show that catching these types of errors can be done efficiently, both at the level of the language and the implementation. Jake Tronge, Howard Pritchard, Jed Brown |
EuroMPI | 1 |
| 2022 | MARS: Malleable Actor-Critic Reinforcement Learning SchedulerabstractIn this paper, we introduce MARS, a new scheduling system for HPC-cloud infrastructures based on a cost-aware, flexible reinforcement learning approach, which serves as an intermediate layer for next generation HPC-cloud resource manager. MARSensembles the pre-trained models from heuristic workloads and decides on the most cost-effective strategy for optimization. A whole workflow application would be split into several optimizable dependent sub-tasks, then based on the predefined resource management plan, a reward will be generated after executing a scheduled task. Lastly, MARSupdates the Deep Neural Network (DNN) model based on the reward. MARSis designed to optimize the existing models through reinforcement mechanisms. MARSadapts to the dynamics of workflow applications, selects the most cost-effective scheduling solution among pre-built scheduling strategies (backfilling, SJF, etc.) and self-learning deep neural network model at run-time. We evaluate MARSwith different real-world workflow traces. MARS can achieve 5%–60% increased performance compared to the state-of-the-art approaches. Betis Baheri, Jake Tronge, Bo Fang 0002, Ang Li 0006, Vipin Chaudhary, Qiang Guan |
IPCCC | 2 |
| 2021 | BEE Orchestrator: Running Complex Scientific Workflows on Multiple SystemsabstractIn this paper, we propose a workflow orchestration system that is able to run workflows on both HPC systems and in the cloud using HPC containers. Most existing workflow orchestration systems are only able to run workflows on one system at a time, and thus may be unable to run workflows that require more resources than what some platforms provide, and may also be unable to handle validation and fault tolerance requirements. Users may have access to a number of different systems, perhaps a mix of HPC systems and private and public clouds, but currently are only able to utilize one system at a time for running complex workflows. Utilizing HPC containers, such as Charliecloud, and a subset of the Common Workflow Language (CWL) for representing workflows, we extend the BEE Orchestration System to allow for possible scheduling of complex workflows across resources. We design and implement a scheduling component and a component for interacting with OpenStack-based HPC clusters and Google Compute Engine clouds to allow for communication between any combination of Cloud and HPC components. We demonstrate how BEE orchestrates workflows across systems, making it possible to run complex scientific applications across all systems that are available to a user. These results also show that BEE will become a viable alternative to other workflow orchestration systems. Jake Tronge, Patricia Grubel, Tim Randles, Quincy Wofford, Rusty Davis, Steven Anaya, Qiang Guan |
HiPC | 1 |
| 2021 | BeeSwarm: Enabling Parallel Scaling Performance Measurement in Continuous Integration for HPC ApplicationsabstractTesting is one of the most important steps in software development–it ensures the quality of software. Continuous Integration (CI) is a widely used testing standard that can report software quality to the developer in a timely manner during development progress. Performance, especially scalability, is another key factor for High Performance Computing (HPC) applications. There are many existing profiling and performance tools for HPC applications, but none of these are integrated into CI tools. In this work, we propose BeeSwarm, an HPC container based parallel scaling performance system that can be easily applied to the current CI test environments. BeeSwarm is mainly designed for HPC application developers who need to monitor how their applications can scale on different compute resources. We demonstrate BeeSwarm using a multi-physics HPC application with Travis CI, GitLab CI and GitHub Actions while using ChameleonCloud and Google Compute Engine as the compute backends. Our results show that BeeSwarm can be used for scalability and performance testing of HPC applications. Jake Tronge, Jieyang Chen, Patricia Grubel, Tim Randles, Rusty Davis, Quincy Wofford, Steven Anaya, Qiang Guan |
ASE | 1 |