EDBT 2026 Demo / reviewers in the wild / expert
Austin R. J. Downey
dblp:217/3558
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-5524-2416ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimized Coding and Parameter Selection for Efficient FPGA Design of Attention MechanismsabstractEfficient utilization of on-chip computational and memory resources, along with optimized high-level synthesis (HLS) coding, is vital to maximize parallelism and minimize latency. This paper demonstrates the HLS algorithms to achieve high utilization of processing elements to enhance parallelism. It also analyzes how various parameters of an attention layer impact latency, employs an efficient tiling technique, and explains the process of selecting an optimized tile size (TS). Ehsan Kabir, Austin R. J. Downey, Jason D. Bakos, David Andrews 0001, Miaoqing Huang |
FCCM | 2 |
| 2025 | N-TORC: Native Tensor Optimizer for Real-Time ConstraintsabstractCompared to overlay-based tensor architectures like VTA or Gemmini, compilers that directly translate machine learning models into a dataflow architecture as HLS code, such as HLS4ML and FINN, generally can achieve lower latency by generating customized matrix-vector multipliers and memory structures tailored to the specific fundamental tensor operations required by each layer. However, this approach has significant drawbacks: the compilation process is highly time-consuming and the resulting deployments have unpredictable area and latency, making it impractical to constrain the latency while simultaneously minimizing area. Currently, no existing methods address this type of optimization. In this paper, we present N-TORC (Native Tensor Optimizer for Real-Time Constraints), a novel approach that utilizes data-driven performance and resource models to optimize individual layers of a dataflow architecture. When combined with model hyperparameter optimization, N-TORC can quickly generate architectures that satisfy latency constraints while simultaneously optimizing for both accuracy and resource cost (i.e. offering a set of optimal trade-offs between cost and accuracy). To demonstrate its effectiveness, we applied this framework to a cyber-physical application, DROPBEAR (Dynamic Reproduction of Projectiles in Ballistic Environments for Advanced Research). N-TORC's HLS4ML performance and resource models achieve higher accuracy than prior efforts, and its Mixed Integer Program (MIP)-based solver generates equivalent solutions to a stochastic search in 1000X less time. Suyash Vardhan Singh, Iftakhar Ahmad, David Andrews 0001, Miaoqing Huang, Austin R. J. Downey, Jason D. Bakos |
FCCM | 5 |
| 2025 | Resource Scheduling for Real-Time Machine Learning
Suyash Vardhan Singh, Iftakhar Ahmad, David Andrews 0001, Miaoqing Huang, Austin R. J. Downey, Jason D. Bakos |
FPGA | 5 |
| 2023 | Accelerating LSTM-Based High-Rate Dynamic System ModelsabstractIn this paper, we evaluate the use of a trained Long Short-Term Memory (LSTM) network as a surrogate for a Euler-Bernoulli beam model, and then we describe and characterize an FPGA-based deployment of the model for use in real-time structural health monitoring applications. The focus of our efforts is the DROPBEAR (Dynamic Reproduction of Projectiles in Ballistic Environments for Advanced Research) dataset, which was generated as a benchmark for the study of real-time structural modeling applications. The purpose of DROPBEAR is to evaluate models that take vibration data as input and give the initial conditions of the cantilever beam on which the measurements were taken as output. DROPBEAR is meant to serve an exemplar for emerging high-rate “active structures” that can be actively controlled with feedback latencies of less than one microsecond. Although the Euler-Bernoulli beam model is a well-known solution to this modeling problem, its computational cost is prohibitive for the time scales of interest. It has been previously shown that a properly structured LSTM network can achieve comparable accuracy with less workload, but achieving sub-microsecond model latency remains a challenge. Our approach is to deploy the LSTM optimized specifically for latency on FPGA. We designed the model using both high-level synthesis (HLS) and hardware description language (HDL). The lowest latency of$1.42\ \mu\mathrm{S}$and the highest throughput of 7.87 Gops/s were achieved on Alveo U55C platform for HDL design. Ehsan Kabir, Daniel Coble, Joud N. Satme, Austin R. J. Downey, Jason D. Bakos, David Andrews 0001, Miaoqing Huang |
FPL | 4 |
| 2023 | Optimal Sampling Methodologies for High-rate Structural TwinningabstractIn high-rate structural health monitoring, it is crucial to quickly and accurately assess the current state of a component under dynamic loads. State information is needed to make informed decisions about timely interventions to prevent damage and extend the structure’s life. In previous studies, a dynamic reproduction of projectiles in ballistic environments (DROPBEAR) testbed was used to evaluate the accuracy of state estimation techniques through dynamic analysis. This paper extends previous research by incorporating the local eigenvalue modification procedure (LEMP) and data fusion techniques to create a more robust state estimate using optimal sampling methodologies. The process of estimating the state involves taking a measured frequency response of the structure, proposing frequency response profiles, and accepting the most similar profile as the new mean for the position estimate distribution. Utilizing LEMP allows for a faster approximation of the proposed model with linear time complexity, making it suitable for 2D or sequential damage cases. The current study focuses on two proposed sampling methodology refinements: distilling the selection of candidate test models from the position distribution and applying a Kalman filter after the distribution update to find the mean. Both refinements were effective in improving the position estimate and the structural state accuracy, as shown by the time response assurance criterion and the signal-to-noise ratio with up to 17% improvement. These two metrics demonstrate the benefits of incorporating data fusion techniques into the high-rate state identification process. Alexander B. Vereen, Emmanuel A. Ogunniyi, Austin R. J. Downey, Erik Blasch, Jason D. Bakos, Jacob Dodson |
FUSION | 3 |
| 2022 | High-Rate Machine Learning for Forecasting Time-Series Signalsabstract"Active structures" are physical structures that incorporate real-time monitoring and control. Examples include active vibration damping or blast mitigation systems. Evaluating physics-based models in real-time is generally not feasible for such systems having high-rate dynamics which require microsecond response times, but data-driven machine-learning-based models can potentially offer a solution. This paper compares the cost and performance of two FPGA-based implementations of real-time, continuously-trained models for forecasting time-series signals with non-stationarities, with one using High-Level Synthesis (HLS) and the other a programmable overlay architecture. The proposed model accepts a uni-variate vibration signal and seeks to forecast future samples to inform high-rate controllers. The proposed forecasting method performs two concurrent neural inference operations. One inference forecasts the state of the signal f samples into the future as a function of the most recent h samples, while the other forecasts the current sample given h samples starting from h+f−1 samples into the past. The first forecast produces the forecast while the second forecast allows the system to calculate the model’s loss and perform an immediate model update before the next sample period. Atiyehsadat Panahi, Ehsan Kabir, Austin R. J. Downey, David Andrews 0001, Miaoqing Huang, Jason D. Bakos |
FCCM | 3 |