Austin Baylis

dblp:214/9803 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0000-7531-2453ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Electronic design automation · 51% Reconfigurable computing and FPGAs · 42% Cloud and datacenter computing · 8%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA accelerator
1.322026
Bridging the Gap: A Module-Context Modeling Methodology for Hyperscale FPGA Applications · FPGA 2026
Scalable Window Generation for the Intel Broadwell+Arria 10 and High-Bandwidth FPGA Systems · FPGA 2018
Electronic design automation › physical design
floorplanning
1.012026
Bridging the Gap: A Module-Context Modeling Methodology for Hyperscale FPGA Applications · FPGA 2026
Electronic design automation
physical design
1.012026
Bridging the Gap: A Module-Context Modeling Methodology for Hyperscale FPGA Applications · FPGA 2026
Reconfigurable computing and FPGAs › reconfigurable computing
FPGA deployment
0.312026
Bridging the Gap: A Module-Context Modeling Methodology for Hyperscale FPGA Applications · FPGA 2026
Cloud and datacenter computing › datacenter architecture
hyperscale datacenter
0.312026
Bridging the Gap: A Module-Context Modeling Methodology for Hyperscale FPGA Applications · FPGA 2026
Machine learning › Deep learning architectures and training › convolutional neural network
CNN inference
0.112018
Scalable Window Generation for the Intel Broadwell+Arria 10 and High-Bandwidth FPGA Systems · FPGA 2018

Methods — techniques the papers use, named apart from their topics

module-context modeling · 1.0pipeline replication · 1.0high-bandwidth memory · 1.0
YearPublicationVenuePosition
2026 Bridging the Gap: A Module-Context Modeling Methodology for Hyperscale FPGA Applications
abstract
Microsoft operates at hyperscale, deploying FPGA accelerators across global datacenters under development realities that diverge from common practices. When full builds take hours or days, or when the complete system doesn't yet exist, developers turn to module isolation to make progress. But existing isolation techniques sacrifice the physical context needed for accurate optimization, so improvements developed in isolation may not survive integration.
Madison N. Emas, Austin Baylis, Greg Stitt
FPGA2
2018 High-Frequency Absorption-FIFO Pipelining for Stratix 10 HyperFlex
abstract
FPGAs often have significantly lower clock frequencies than microprocessors and GPUs, due largely to propagation delays incurred by the reconfigurable interconnect. The Stratix 10 HyperFlex architecture reduces this problem by embedding numerous registers throughout the routing resources. However, such Hyper-Registers do not support back-pressure (i.e., pipeline stalls) that is commonly used in FPGA pipelines. In this paper, we present and evaluate pipeline transformations using absorption FIFOs, which avoid back-pressure limitations to enable numerous pipelines to benefit from HyperFlex, while also eliminating potentially expensive stall penalties incurred by existing techniques. We demonstrate that these transformations not only enable significant clock improvements on Stratix 10, but also for devices without HyperFlex, potentially making absorption FIFOs a better high-frequency strategy for any FPGA.
Madison N. Emas, Austin Baylis, Greg Stitt
FCCM2
2018 Scalable Window Generation for the Intel Broadwell+Arria 10 and High-Bandwidth FPGA Systems
abstract
Emerging FPGA systems are providing higher external memory bandwidth to compete with GPU performance. However, because FPGAs often achieve parallelism through deep pipelines, traditional FPGA design strategies do not necessarily scale well to large amounts of replicated pipelines that can take advantage of higher bandwidth. We show that sliding-window applications, an important subset of digital signal processing, demonstrate this scalability problem. We introduce a window generator architecture that enables replication to over 330 GB/s, which is an 8.7x improvement over previous work. We evaluate the window generator on the Intel Broadwell+Arria10 system for 2D convolution and show that for traditional convolution (one filter per image), our approach outperforms a 12-core Xeon Broadwell E5 by 81x and a high-end Nvidia P6000 GPU by an order of magnitude for most input sizes, while improving energy by 15.7x. For convolutional neural nets (CNNs), we show that although the GPU and Xeon typically outperform existing FPGA systems, projected performances of the window generator running on FPGAs with sufficient bandwidth can outperform high-end GPUs for many common CNN parameters.
Greg Stitt, Abhay Gupta, Madison N. Emas, David Wilson 0004, Austin Baylis
FPGA5