Aaron R. Young

dblp:228/1509 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-5448-4667ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 9 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PowerMappeR: Power-Optimized Mapping of SNNs onto ReRAM Crossbars coupled via Packet-Switched NoCs
Devin Pohl, Kazi Asifuzzaman, Aaron R. Young, Narasinga Rao Miniskar, Jeffrey S. Vetter
IPDPS3
2026 Accl++ : A high-productivity programming language for performance and code portability on heterogeneous systems
abstract
This work describes the Accl++ programming language for heterogeneous computing. Accl++ is embedded in the C++ language and implemented as a C++ library. The language allows for describing device code and execution libraries and includes primitives for runtime compilation (RTC), thereby enabling code portability across disparate devices. Here, we demonstrate Accl++’s capability by coding different benchmarks from different domains. This work also describes an analysis of the overheads introduced by the Accl++ RTC support as well as the performance of the generated code for the Accl++ kernels. Accl++ improves heterogeneous code portability without incurring high levels of overhead. In two different heterogeneous systems—one composed of two 32-core AMD EPYC 7513 CPUs and two NVIDIA A100 GPUs and the other composed of two 12-core AMD EPYC 7272 CPUs and two AMD MI100 GPUs—Accl++ enables the execution of the same binary application with observed RTC overheads in the range of 3%–10% of the total kernel execution time, resulting in performance levels similar to those of native CUDA/HIP/OpenCL code.
Marc González 0001, Pedro Valero-Lara, Mohammad Alaul Haque Monil, Seyong Lee, Beau Johnston, Aaron R. Young, Narasinga Rao Miniskar, Keita Teranishi, Jeffrey S. Vetter
Future Gener. Comput. Syst.6
2025 Mapping Spiking Neural Networks to Heterogeneous Crossbar Architectures using Integer Linear Programming
abstract
Advances in novel hardware devices and architectures allow Spiking Neural Network (SNN) evaluation using ultra-low power, mixed-signal, memristor crossbar arrays. As individual network sizes quickly scale beyond the dimensional capabilities of single crossbars, networks must be mapped onto multiple crossbars. Crossbar sizes within modern Memristor Crossbar Architectures (MCAs) are determined predominately not by device technology but by network topology; more, smaller crossbars consume less area thanks to the high structural sparsity found in larger, brain-inspired SNNs. Motivated by continuing increases in SNN sparsity due to improvements in training methods, we propose utilizing heterogeneous crossbar sizes to further reduce area consumption. This approach was previously unachievable as prior compiler studies only explored solutions targeting homogeneous MCAs. Our work improves on the state-of-the-art by providing Integer Linear Programming (ILP) formulations supporting arbitrarily heterogeneous architectures. By modeling axonal interactions between neurons, our methods produce better mappings while removing inhibitive a priori knowledge requirements. We first show a 16.7-27.6% reduction in area consumption for square-crossbar homogeneous architectures. Then, we demonstrate 66.9-72.7% further reduction when using a reasonable configuration of heterogeneous crossbar dimensions. Next, we present a new optimization formulation capable of minimizing the number of inter-crossbar routes. When applied to solutions already near-optimal in area, an 11.9-26.4% routing reduction is observed without impacting area consumption. Finally, we present a profile-guided optimization capable of minimizing the number of runtime spikes between crossbars. Compared to the best-area-then-route optimized solutions, we observe a further 0.5-14.8% inter-crossbar spike reduction while requiring 1–3 orders of magnitude less solver time.
Devin Pohl, Aaron R. Young, Kazi Asifuzzaman, Narasinga Rao Miniskar, Jeffrey S. Vetter
DATE2
2025 ReSpike: A Co-Design Framework for Evaluating SNNs on ReRAM-Based Neuromorphic Processors
Kazi Asifuzzaman, Aaron R. Young, Prasanna Date, Shruti R. Kulkarni, Narasinga Rao Miniskar, Matthew J. Marinella, Jeffrey S. Vetter
Euro-Par (2)2
2025 IRIS-MASH: Efficient Multi-device Asynchronous Multi-Stream Heterogeneous Computing
abstract
In the rapidly evolving field of high-performance computing (HPC), effectively leveraging heterogeneous devices through asynchronous task programming is paramount. This paper presents a robust asynchronous task programming model tailored for a multi-device, multi-stream execution environment that incorporates a diverse array of heterogeneous computing units, including GPUs from various vendors and other accelerators. Current state-of-the-art task programming models provide methodologies to support asynchronous task executions, but they typically handle homogeneous devices using native programming languages, while support for heterogeneous devices is limited to frameworks like OpenCL. This gap presents significant challenges in abstracting heterogeneous devices to harness their true asynchronous capabilities effectively using their native programming languages. By implementing asynchronous task execution, our model significantly boosts the performance of tiled algorithm task graphs through overlapping data transfers with computation and enabling the simultaneous execution of multiple kernels. We integrate this approach into a heterogeneous Intelligent Runtime System (IRIS) and assess its performance using a suite of tiled algorithm benchmarks from the heterogeneous math kernels library (MatRIS) based on IRIS. Experimental results demonstrate a performance improvement ranging from 1.6 × to 2 × over IRIS without asynchronous support, and a notable 22% performance enhancement compared to established runtime systems such as StarPU and PaRSEC. This approach significantly improves computation efficiency of HPC workflows and provides a solid base for future exploration and development in the area of asynchronous task programming in heterogeneous systems.
Narasinga Rao Miniskar, Aaron R. Young, Mohammad Alaul Haque Monil, Kazi Asifuzzaman, Beau Johnston, Keita Teranishi, Jeffrey S. Vetter
ICPP2
2025 Exploring Spiking Neural Networks for Binary Classification in Multivariate Time Series at the Edge
abstract
We present a general framework for training spiking neural networks (SNNs) to perform binary classification on multivariate time series, with a focus on step-wise prediction and high precision at low false alarm rates. The approach uses the Evolutionary Optimization of Neuromorphic Systems (EONS) algorithm to evolve sparse, stateful SNNs by jointly optimizing their architectures and parameters. Inputs are encoded into spike trains, and predictions are made by thresholding a single output neuron’s spike counts. We also incorporate simple voting ensemble methods to improve performance and robustness.To evaluate the framework, we apply it with application-specific optimizations to the task of detecting low signal-to-noise ratio radioactive sources in gamma-ray spectral data. The resulting SNNs, with as few as 49 neurons and 66 synapses, achieve a 51.8% true positive rate (TPR) at a false alarm rate of 1/hr, outperforming PCA (42.7%) and deep learning (49.8%) baselines. A three-model any-vote ensemble increases TPR to 67.1% at the same false alarm rate. Hardware deployment on the μCaspian neuromorphic platform demonstrates 2 mW power consumption and 20.2 ms inference latency.We also demonstrate generalizability by applying the same framework, without domain-specific modification, to seizure detection in EEG recordings. An ensemble achieves 95% TPR with a 16% false positive rate, comparable to recent deep learning approaches with significant reduction in parameter count.
James Ghawaly, Andrew D. Nicholson, Catherine D. Schuman, Dalton Diez, Aaron R. Young, Brett Witherspoon
IJCNN5
2025 ChatHPC: Building the Foundations for a Productive and Trustworthy AI-Assisted HPC Ecosystem
abstract
ChatHPC democratizes large language models for the high-performance computing (HPC) community by providing the infrastructure, ecosystem, and knowledge needed to apply modern generative AI technologies to rapidly create specific capabilities for critical HPC components while using relatively modest computational resources. Our divide-and-conquer approach focuses on creating a collection of reliable, highly specialized, and optimized AI assistants for HPC based on the cost-effective and fast Code Llama fine-tuning processes and expert supervision. We target major components of the HPC software stack, including programming models, runtimes, I/O, tooling, and math libraries. Thanks to AI, ChatHPC provides a more productive HPC ecosystem by boosting important tasks related to portability, parallelization, optimization, scalability, and instrumentation, among others. With relatively small datasets (on the order of KB), the AI assistants, which are created in a few minutes by using one node with two NVIDIA H100 GPUs and the ChatHPC library, can create new capabilities with Meta’s 7-billion parameter Code Llama base model to produce high-quality software with a level of trustworthiness of up to 90% higher than the 1.8-trillion parameter OpenAI ChatGPT-4o model for critical programming tasks in the HPC software stack.
Pedro Valero-Lara, Aaron R. Young, Jeffrey S. Vetter, Zheming Jin, Swaroop Pophale, Mohammad Alaul Haque Monil, Keita Teranishi, William F. Godoy
SC2
2024 Event-Driven Sensing and Embedded Neuromorphic Platforms for Gamma Radiation Monitoring
abstract
This work will present an embedded neuromorphic platform developed for long-term unintended gamma radiation monitoring. We describe the hardware architecture and supporting software developed to demonstrate neuromorphic computing for applications where ultra-low power and always-on sensing are required. This is followed by a discussion of our current work on an improved platform that integrates both event-driven vision and gamma-ray spectroscopy sensors for nuclear safeguards applications. Finally, future research directions toward event-driven sampling techniques to integrate analog-to-information reduction into the sensing electronics are proposed.
Brett Witherspoon, Aaron R. Young
ACM Great Lakes Symposium on VLSI2
2023 A 3D Implementation of Convolutional Neural Network for Fast Inference
abstract
Low latency inference has many applications in edge machine learning. In this paper, we present a run-time configurable convolutional neural network (CNN) inference ASIC design for low-latency edge machine learning. By implementing a 5-stage pipelined CNN inference model in a 3D ASIC technology, we demonstrate that the model distributed on two dies utilizing face-to-face (F2F) 3D integration achieves superior performance. Our experimental results show that the design based on 3D integration achieves 43% better energy-delay product when compared to the traditional 2D technology.
Narasinga Rao Miniskar, Pruek Vanna-Iampikul, Aaron R. Young, Sung Kyu Lim, Frank Liu 0001, Jieun Yoo, Corrinne Mills, Farah Fahim, Jeffrey S. Vetter
ISCAS3
2022 Ultra Low Latency Machine Learning for Scientific Edge Applications
abstract
In this paper, we present an FPGA design of an extremely low latency scientific machine learning application at the edge. Real-time prediction of errant high-energy particle beams at scientific facilities such as Spallation Neutron Source (SNS) is crucial to avoid damages to the equipment. Machine learning techniques are becoming increasingly effective to detect subtle signatures of the errant beams in the noisy sensor signals. However, to minimize potential damage done by errant beam, real-time errant beam detection has to be completed with extremely low latency, usually less than 1 microsecond. By stream processing the input features and employing out-of-order execution of decision nodes among the decision trees, we demonstrate that our highly efficient FPGA implementation can achieve 60 nanoseconds of computing latency for complex random forest models with 10,000 input features.
Narasinga Rao Miniskar, Aaron R. Young, Frank Liu 0001, Willem Blokland, Anthony M. Cabrera, Jeffrey S. Vetter
FPL2
2020 Scaled-up Neuromorphic Array Communications Controller (SNACC) for Large-scale Neural Networks
abstract
Neuromorphic computing is one promising post-Moore's law era technology, which takes inspiration from biological brains to perform computing tasks. The human brain contains billions of neurons with trillions of synapses and as neuromorphic hardware systems scale to larger and larger sizes, the communication system used to transfer information between neuromorphic elements and traditional computers must scale to keep up. In prior work, we describe the use of a separate neuromorphic array communications controller to support low-latency, high-throughput communication between our neuromorphic systems and a traditional computer. In this work, the neuromorphic array communications controller is used to support the scaling of a neuromorphic development system which uses multiple neuromorphic processors arranged in a two-dimensional array. The neuromorphic array communications controller, along with scalable local connections, is used to create a scalable neuromorphic platform to enable the development and testing of large neuromorphic network arrays.
Aaron R. Young, Adam Z. Foshie, Mark E. Dean, James S. Plank, Garrett S. Rose, J. Parker Mitchell, Catherine D. Schuman
IJCNN1
2018 Understanding Selection And Diversity For Evolution Of Spiking Recurrent Neural Networks
abstract
Evolutionary optimization or genetic algorithms have been used to optimize a variety of neural network types, including spiking recurrent neural networks, and are attractive for many reasons. However, a key impediment to their widespread use is the potential for slow training times and failure to converge to a good fitness value in a reasonable amount of time. In this work, we evaluate the effect of different selection algorithms on the performance of an evolutionary optimization method for designing spiking recurrent neural networks, including those that are meant to be deployed in a neuromorphic system. We propose a selection approach that utilizes a richer understanding of the fitness of an individual network to inform the selection process and to promote diversity in the population. We show that including this feature can provide a significant increase in performance over utilizing a standard selection approach.
Catherine D. Schuman, Grant Bruer, Aaron R. Young, Mark E. Dean, James S. Plank
IJCNN3
2018 Neuromorphic Array Communications Controller to Support Large-Scale Neural Networks
abstract
Neuromorphic computing is one promising post-Moore's law era technology. In order to develop and use neuromorphic systems, traditional von Neumann-based computers must be able to communicate with neuromorphic hardware to support functionality such as monitoring the state of the network, optimizing the array to better perform the task, and input/output data processing. In this paper, we describe our use of a separate neuromorphic array communications controller to support highthroughput, low-latency communication between a traditional computer and our implementations of neuromorphic systems. The goal of the communications controller is to provide enough performance to facilitate the desired interaction between the systems and to enable scaling of the neuromorphic systems to larger sizes.
Aaron R. Young, Mark E. Dean, James S. Plank, Garrett S. Rose, Catherine D. Schuman
IJCNN1