Jeffrey Young 0001

dblp:198/8661 · also Jeff Young 0001, Jeffrey S. Young 0001, Jeffrey Scott Young 0001 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0001-9841-4057ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 ASDF: A Compiler for Qwerty, a Basis-Oriented Quantum Programming Language
abstract
Qwerty is a high-level quantum programming language built on bases and functions rather than circuits. This new paradigm introduces new challenges in compilation, namely synthesizing circuits from basis translations and automatically specializing adjoint or predicated forms of functions. This paper presents ASDF, an open-source compiler for Qwerty that answers these challenges in compiling basis-oriented languages. Enabled with a novel high-level quantum IR implemented in the MLIR framework, our compiler produces OpenQASM 3 or QIR for either simulation or execution on hardware. Our compiler is evaluated by comparing the fault-tolerant resource requirements of generated circuits with other compilers, finding that ASDF produces circuits with comparable cost to prior circuit-oriented compilers.
Austin J. Adams, Sharjeel Khan, Arjun S. Bhamra, Ryan R. Abusaada, Anthony M. Cabrera, Cameron C. Hoechst, Travis S. Humble, Jeffrey Young 0001, Thomas M. Conte
CGO8
2025 A Blueprint for Q-CS1, an Introductory Quantum Programming Course
abstract
Despite the need to build a quantum workforce, current courses that introduce quantum programming are rooted in quantum notation that students may find intimidating. We propose Q-CS1, a quantum equivalent of CS1 that begins with hands-on quantum programming. Q-CS1 is enabled by the Qwerty quantum programming language, which allows for reasoning about qubit behavior without physics notation or quantum circuits. An outline of Q-CS1 is provided along with plans for assessing its effectiveness.
Austin J. Adams, Rodrigo Borela, Jeffrey Young 0001, Thomas M. Conte
SIGCSE (2)3
2024 Asynchronous Distributed-Memory Parallel Algorithms for Influence Maximization
abstract
Influence maximization (IM) is the problem of finding the k most influential nodes in a graph. We propose distributed-memory parallel algorithms for the two main kernels of a state-of-the-art implementation of one IM algorithm, influence maximization via martingales (IMM). The baseline relies on a bulk-synchronous parallel approach and uses replication to reduce communication and achieve approximate load balance, at the cost of synchronization and high memory requirements. By contrast, our method fully distributes the data, thereby improving memory scalability, and uses fine-grained asynchronous parallelism to improve network utilization and the cost of doing more communication. We show our design and implementation can achieve up to $29.6 \times$ speedup over the MPI-based state-of-the-art on synthetic and real-world network graphs. Moreover, ours is the first implementation that can run IMM to find influencers in the ‘twitter’ graph (41M nodes and 1.4B edges) in 200 seconds using 8 K CPU cores of NERSC Perlmutter supercomputer.
Shubhendra Pal Singhal, Souvadra Hati, Jeffrey Young 0001, Vivek Sarkar, Akihiro Hayashi, Richard W. Vuduc
SC3
2024 CuPBoP: Making CUDA a Portable Language
abstract
CUDA is designed specifically for NVIDIA GPUs and is not compatible with non-NVIDIA devices. Enabling CUDA execution on alternative backends could greatly benefit the hardware community by fostering a more diverse software ecosystem. To address the need for portability, our objective is to develop a framework that meets key requirements, such as extensive coverage, comprehensive end-to-end support, superior performance, and hardware scalability. Existing solutions that translate CUDA source code into other high-level languages, however, fall short of these goals. In contrast to these source-to-source approaches, we present a novel framework, CuPBoP , which treats CUDA as a portable language in its own right. Compared to two commercial source-to-source solutions, CuPBoP offers a broader coverage and superior performance for the CUDA-to-CPU migration. Additionally, we evaluate the performance of CuPBoP against manually optimized CPU programs, highlighting the differences between CPU programs derived from CUDA and those that are manually optimized. Furthermore, we demonstrate the hardware scalability of CuPBoP by showcasing its successful migration of CUDA to AMD GPUs. To promote further research in this field, we have released CuPBoP as an open-source resource.
Ruobing Han, Jun Chen 0038, Bhanu Garg, Xule Zhou, John Lu, Jeffrey Young 0001, Jaewoong Sim, Hyesoon Kim
ACM Trans. Design Autom. Electr. Syst.6
2023 CuPBoP: A Framework to Make CUDA Portable
abstract
CUDA, as one of the most popular choices for GPU programming, can be executed only on NVIDIA GPUs. To execute CUDA on non-NVIDIA devices, researchers have proposed to translate CUDA to other programming languages. However, this approach cannot achieve high coverage due to the challenges in source-to-source translation.
Ruobing Han, Jun Chen 0038, Bhanu Garg, Jeffrey Young 0001, Jaewoong Sim, Hyesoon Kim
PPoPP4
2023 Towards Safe HPC: Productivity and Performance via Rust Interfaces for a Distributed C++ Actors Library (Work in Progress)
abstract
In this work-in-progress research paper, we make the case for using Rust to develop applications in the High Performance Computing (HPC) domain which is critically dependent on native C/C++ libraries. This work explores one example of Safe HPC via the design of a Rust interface to an existing distributed C++ Actors library. This existing library has been shown to deliver high performance to C++ developers of irregular Partitioned Global Address Space (PGAS) applications.
John Parrish, Nicole Wren, Tsz Hang Kiang, Akihiro Hayashi, Jeffrey Young 0001, Vivek Sarkar
MPLR5
2023 HIPLZ: Enabling performance portability for exascale systems
abstract
Summary While heterogeneous computing has emerged as a dominant trend in current and future High‐Performance Computing (HPC) systems, it is also widely recognized that this shift has led to increased software complexity due to a proliferation of programming systems for different heterogeneous processors. One such example is the Heterogeneous‐Compute Interface for Portability from AMD (HIP ), which is composed of a C Runtime API and C++ Kernel Language. Many HPC applications will likely use HIP on future exascale systems (e.g., Frontier and El Capitan), but HIP currently only targets AMD and NVIDIA processors. This limitation creates challenges for users who would also like to run their applications on exascale systems based on other architectures (e.g., Aurora, which is based on Intel hardware) that are currently not targeted by HIP . In this paper, we introduce the design and implementation of HIPLZ , a compiler and runtime system that uses the Intel Level Zero API to support HIP on Intel GPU architectures. We discuss the design of HIPLZ , derived from HIPCL (an implementation of HIP on top of OpenCL ), and portability issues that occur from using the Level Zero runtime as a backend. We evaluate our implementation by running several performance benchmarks and mini‐apps written in HIP on Intel architectures using HIPLZ . Our results show that this approach provides competitive performance relative to Intel's OpenCL implementations on Intel Gen9 and UHD Graphics 770 GPUs, while providing good coverage of features needed by HPC applications. Overall, this approach is a promising demonstration of enabling performance portability for exascale systems.
Jisheng Zhao, Colleen Bertoni, Jeffrey Young 0001, Kevin Harms, Vivek Sarkar, Brice Videau
Concurr. Comput. Pract. Exp.3
2022 Accelerating Graphic Rendering on Programmable RISC-V GPUs
abstract
Graphics rendering remains one of the most compute-intensive and memory-bound applications of GPUs and has been driving their push for performance and energy efficiency since its inception. Early GPU architectures focused only on accelerating graphics rendering and implemented dedicated a fixed-function rendering units. Today’s GPUs have become more programmable to address the complexity and diversity of modern graphics workloads while still accelerating several components of the graphics pipeline in fixed-function hardware.Generalizing the GPU microarchitecture and implement some of its graphics hardware blocks in software can save area that can be used to expand the generic pipeline, especially in mobile systems-on-chips environments where power and area is scarce.In this work, we propose a RISC-V-based hybrid GPU architecture that accelerates the graphics pipeline without paying the cost of a full hardware graphics pipeline. We evaluated the design on an Altera Arria 10 FPGA running at 200 MHz.
Blaise-Pascal Tine, Varun Saxena, Santosh Srivatsan, Joshua R. Simpson, Fadi Alzammar, Liam Cooper, Sam Jijina, Swetha Rajagoplan, Tejaswini Anand Kumar, Jeffrey Young 0001, Hyesoon Kim
HCS10
2022 ParaGraph: An application-simulator interface and toolkit for hardware-software co-design
abstract
ParaGraph is an open-source toolkit for use in co-designing hardware and software for supercomputer-scale systems. It bridges an infrastructure gap between an application target and existing high-fidelity computer-network simulators. The first component of ParaGraph is a high-level graph representation of a parallel program, which a) faithfully represents parallelism and communication, b) can be extracted automatically from a compiler, and c) is “tuned” for use with network simulators. The second is a runtime that can emulate the representation’s dynamic execution for a simulator. User-extensible mechanisms are available for modeling on-node performance and transforming high-level communication into operations that backend simulators understand. Case studies include deep learning workloads that are extracted automatically from programs written in JAX and TensorFlow and interfaced with several event-driven network simulators. These studies show how system designers can use ParaGraph to build flexible end-to-end software-hardware co-design workflows to tweak communication libraries, find future hardware bottlenecks, and validate simulations with traces.
Mikhail Isaev, Nic McDonald, Jeffrey Young 0001, Richard W. Vuduc
ICPP3
2022 "Smarter" NICs for faster molecular dynamics: a case study
abstract
This work evaluates the benefits of using a “smart” network interface card (SmartNIC) as a compute accelerator for the example of the MiniMD molecular dynamics proxy application. The accelerator is NVIDIA's BlueField-2 card, which includes an 8-core Arm processor along with a small amount of DRAM and storage. We test the networking and data movement performance of these cards compared to a standard Intel server host using microbenchmarks and MiniMD. In MiniMD, we identify two distinct classes of computation, namely core computation and maintenance computation, which are executed in sequence. We restructure the algorithm and code to weaken this dependence and increase task parallelism, thereby making it possible to increase utilization of the BlueField-2 concurrently with the host. We evaluate our implementation on a cluster consisting of 16 dual-socket Intel Broadwell host nodes with one BlueField-2 per host-node. Our results show that while the overall compute performance of BlueField-2 is limited, using them with a modified MiniMD algorithm allows for up to 20% speedup over the host CPU baseline with no loss in simulation accuracy.
Sara Karamati, Clay Hughes, Karl S. Hemmert, Ryan E. Grant, Whit Schonbein, Scott Levy, Thomas M. Conte, Jeffrey Young 0001, Richard W. Vuduc
IPDPS8
2021 Online model swapping for architectural simulation
abstract
As systems and applications grow more complex, detailed computer architecture simulation takes an ever increasing amount of time. Longer simulation times result in slower design iterations which then force architects to use simpler models, such as spreadsheets, when they want to iterate quickly on a design. Simple models are not easy to work with though, as architects must rely on intuition to choose representative models, and the path from the simple models to a detailed hardware simulation is not always clear.
Patrick Lavin, Jeffrey Young 0001, Richard W. Vuduc, Jonathan Beard
CF2
2020 RISC-V FPGA Platform Toward ROS-Based Robotics Application
abstract
RISC-V is free and open standard instruction set architecture following reduced instruction set computer principle. Because of its openness and scalability, RISC-V has been adapted not only for embedded CPUs such as mobile and IoT market, but also for heavy-workload CPUs such as the data center or super computing field. On top of it, Robotics is also a good application of RISC-V because security and reliability become crucial issues of robotics system. These problems could be solved by enthusiastic open source community members as they have shown on open source operating system. However, running RISC-V on local FPGA becomes harder than before because now RISC-V foundation are focusing on cloud-based FPGA environment. We have experienced that recently released OS and toolchains for RISC-V are not working well on the previous CPU image for local FPGA. In this paper we design the local FPGA platform for RISC-V processor and run the robotics application on mainstream Robot Operating System on top of the RISC-V processor. This platform allow us to explore the architecture space of RISC-V CPU for robotics application, and get the insight of the RISC-V CPU architecture for optimal performance and the secure system.
Hanning Chen, Jeffrey Young 0001, Hyesoon Kim
FPL3
2019 A microbenchmark characterization of the Emu chick
Jeffrey Young 0001, Eric R. Hein, Srinivas Eswar, Patrick Lavin, Jiajia Li 0001, E. Jason Riedy, Richard W. Vuduc, Thomas M. Conte
Parallel Comput.1
2018 An Energy-Efficient Single-Source Shortest Path Algorithm
abstract
We present a novel strategy to control the energy-efficiency of an algorithm from software, which is to make the degree of parallelism dynamically and automatically tunable. The specific algorithm is a variation of delta-stepping for computing a single-source shortest path (SSSP); its available parallelism is highly irregular and strongly input-dependent. Informed by an analysis of these runtime characteristics, we propose a software-based controller that uses online learning techniques to tune parallelism to meet a given target, thereby improving the average available parallelism while reducing its variability. We show experimentally the efficacy of our self-tuning algorithm in managing tradeoffs among performance and power. Our experimental apparatus is based on the SSSP implementation available in the Gunrock GPU library running on an embedded CPU+GPU, whose hardware has GPU core and memory frequency knobs.
Sara Karamati, Jeffrey Young 0001, Richard W. Vuduc
IPDPS2
2018 Performance Implications of NoCs on 3D-Stacked Memories: Insights from the Hybrid Memory Cube
abstract
Three-dimensional (3D)-stacked memories, such as the Hybrid Memory Cube (HMC), provide a promising solution for overcoming the bandwidth wall between processors and memory by integrating memory and logic dies in a single stack. Such memories also utilize a network-on-chip (NoC) to connect their internal structural elements and to enable scalability. This novel usage of NoCs enables numerous benefits such as high bandwidth and memory-level parallelism and creates future possibilities for efficient processing-in-memory techniques. However, the implications of such NoC integration on the performance characteristics of 3D-stacked memories in terms of memory access latency and bandwidth have not been fully explored. This paper addresses this knowledge gap (i) by characterizing an HMC prototype using Micron's AC-510 accelerator board and by revealing its access latency and bandwidth behaviors; and (ii) by investigating the implications of such behaviors on system- and software-level designs. Compared to traditional DDR-based memories, our examinations reveal the performance impacts of NoCs for current and future 3D-stacked memories and demonstrate how the packet-based protocol, internal queuing characteristics, traffic conditions, and other unique features of the HMC affects the performance of applications.
Ramyad Hadidi, Bahar Asgari, Jeffrey Young 0001, Burhan Ahmad Mudassar, Kartikay Garg, Tushar Krishna, Hyesoon Kim
ISPASS3
2016 Landrush: Rethinking In-Situ Analysis for GPGPU Workflows
abstract
In-situ analysis on the output data of scientific simulations has been made necessary by ever-growing output data volumes and increasing costs of data movement as supercomputing is moving towards exascale. With hardware accelerators like GPUs becoming increasingly common in high end machines, new opportunities arise to co-locate scientific simulations and online analysis performed on the scientific data generated by the simulations. However, the asynchronous nature of GPGPU programming models and the limited context-switching capabilities on the GPU pose challenges to co-locating the scientific simulation and analysis on the same GPU. This paper dives deeper into these challenges to understand how best to co-locate analysis with scientific simulations on the GPUs in HPC clusters. Specifically, our 'Landrush' approach to GPU sharing proposes a solution that utilizes idle cycles on the GPU to provide an improved time-to-answer, that is, the total time to run the scientific simulation and analysis of the generated data. Landrush is demonstrated with experimental results obtained from leadership high-end applications on ORNL's Titan supercomputer, which show that (i) GPU-based scientific simulations have varying degrees of idle cycles to afford useful analysis task co-location, and (ii) the inability to context switch on the GPU at instruction granularity can be overcome by careful control of the analysis kernel launches and software-controlled early completion of analysis kernel executions. Results show that Landrush is superior in terms of time-to-answer compared to serially running simulations followed by analysis or by relying on the GPU driver and hardwired thread dispatcher to run analysis concurrently on a single GPU.
Anshuman Goswami, Yuan Tian 0004, Karsten Schwan, Fang Zheng 0003, Jeffrey Young 0001, Matthew Wolf, Greg Eisenhauer, Scott Klasky
CCGrid5
2016 GraphIn: An Online High Performance Incremental Graph Processing Framework
Dipanjan Sengupta, Narayanan Sundaram, Theodore L. Willke, Jeffrey Young 0001, Matthew Wolf, Karsten Schwan
Euro-Par5
2013 Oncilla: A GAS runtime for efficient resource allocation and data movement in accelerated clusters
abstract
Accelerated and in-core implementations of Big Data applications typically require large amounts of host and accelerator memory as well as efficient mechanisms for transferring data to and from accelerators in heterogeneous clusters. Scheduling for heterogeneous CPU and GPU clusters has been investigated in depth in the high-performance computing (HPC) and cloud computing arenas, but there has been less emphasis on the management of cluster resource that is required to schedule applications across multiple nodes and devices. Previous approaches to address this resource management problem have focused on either using low-performance software layers or on adapting complex data movement techniques from the HPC arena, which reduces performance and creates barriers for migrating applications to new heterogeneous cluster architectures. This work proposes a new system architecture for cluster resource allocation and data movement built around the concept of managed Global Address Spaces (GAS), or dynamically aggregated memory regions that span multiple nodes.We propose a software layer called Oncilla that uses a simple runtime and API to take advantage of non-coherent hardware support for GAS. The Oncilla runtime is evaluated using two different high-performance networks for microkernels representative of the TPC-H data warehousing benchmark, and this runtime enables a reduction in runtime of up to 81%, on average, when compared with standard disk-based data storage techniques. The use of the Oncilla API is also evaluated for a simple breadth-first search (BFS) benchmark to demonstrate how existing applications can incorporate support for managed GAS.
Jeffrey Young 0001, Se Hoon Shon, Sudhakar Yalamanchili, Alex Merritt, Karsten Schwan, Holger Fröning
CLUSTER1
2006 Poster reception - Parallel I/O advancements in air quality modeling systems
abstract
This poster presents an overview of and the performance results of recent I/O advancements in the parallel CMAQ framework. These optimizations were developed as part of a collaboration between the EPA and Sandia National Laboratories. netCDF provides a portable file format and an easily understood API, but it does not support concurrent writes by multiple processes. In a cluster environment, this leaves two basic options: create a file per process or funnel all the data through a single node. Neither of these options are optimal. Using a slightly modified API, parallel-netCDF (pnetCDF) enables high performance parallel I/O using the MPI-IO collective I/O optimizations while maintaining the netCDF file format. We have created a thin wrapper around pnetCDF that makes it simple for users to enable the new parallel I/O features at compile time. Using parallel I/O has improved the write performance of the CMAQ air-quality modeling code by up to 48%.
Todd Kordenbrock, Ron A. Oldfield, Jeffrey Young 0001
SC3