Timothy Martin

dblp:49/2352 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
3since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2024 A High-Performance Routing Engine for Large-Scale FPGAs
abstract
Routing is the most time-consuming stage in the Field Programmable Gate Array (FPGA) design workflow. We propose a parallel routing technology, based on the Pathfinder algorithm, that enhances parallelism by dividing the search into two phases: one that tolerates overlaps and one that does not. Additional performance optimizations include an improved cost schedule, pruning the routing-resource graph, and selecting efficient data structures for modern CPUs. Evaluated using both the 2023 MLCAD and 2024 FPGA Routing Contest benchmarks, our router achieves average speedups of $6.2 \times$ and $5.2 \times$ compared to RWRoute and Vivado 2023.2, respectively.
Timothy Martin, Dani Maarouf, Gary William Grewal, Shawki Areibi
FPL1
2022 Guiding FPGA Detailed Placement via Reinforcement Learning
abstract
Detailed Placement (DP) is an important, but time-consuming, optimization step within the Field Programmable Gate Array (FPGA) design flow. Given a global placement, DP seeks to refine the global placement to improve the success of the subsequent routing step. In this paper, we show how Reinforcement Learning (RL) can be used to significantly reduce DP runtimes while maintaining Quality-of-Result (QoR). We develop 3 different RL models based on Tabular Q-Learning, Deep Q-Learning, and Actor-Critic. These models are evaluated by integrating them into GPlace3.0 – a state-of-the-art analytic FPGA placement tool – and tested using the 12 ISPD contest benchmarks. Our results show the models achieve total runtime improvements between 2x to 3.5x and similar QoR compared to GPlace3.0’s algorithmic-based detailed placer.
P. Esmaeili, Timothy Martin, Shawki Areibi, Gary William Grewal
VLSI-SoC2
2021 A Deep Learning Framework to Predict Routability for FPGA Circuit Placement
abstract
The ability to accurately and efficiently estimate the routability of a circuit based on its placement is one of the most challenging and difficult tasks in the Field Programmable Gate Array (FPGA) flow. In this article, we present a novel, deep learning framework based on a Convolutional Neural Network (CNN) model for predicting the routability of a placement. Since the performance of the CNN model is strongly dependent on the hyper-parameters selected for the model, we perform an exhaustive parameter tuning that significantly improves the model’s performance and we also avoid overfitting the model. We also incorporate the deep learning model into a state-of-the-art placement tool and show how the model can be used to (1) avoid costly, but futile, place-and-route iterations, and (2) improve the placer’s ability to produce routable placements for hard-to-route circuits using feedback based on routability estimates generated by the proposed model. The model is trained and evaluated using over 26K placement images derived from 372 benchmarks supplied by Xilinx Inc. We also explore several opportunities to further improve the reliability of the predictions made by the proposed DLRoute technique by splitting the model into two separate deep learning models for (a) global and (b) detailed placement during the optimization process. Experimental results show that the proposed framework achieves a routability prediction accuracy of 97% while exhibiting runtimes of only a few milliseconds.
Abeer Alhyari, Hannah Szentimrey, Ahmed Elshamli, Timothy Martin, Gary William Grewal, Shawki Areibi
ACM Trans. Reconfigurable Technol. Syst.4
2020 A Deep-Learning Framework for Predicting Congestion During FPGA Placement
abstract
The ability to quickly and accurately predict congestion has emerged as one of the most critical problems during placement. In this paper, we present DLCong, a deep learning congestion-estimation framework based on a convolutional encoder-decoder. Experimental results show that compared to MLCong, a state-of-the-art machine-learning based congestion-estimation model, DLCong achieves an almost 9% improvement in congestion accuracy, while exhibiting inference times of a few milliseconds. Moreover, the accuracy of DLCong scales better with increasing congestion compared to MLCong.
Dani Maarouf, Ahmed Elshamli, Timothy Martin, Gary William Grewal, Shawki Areibi
FPL3
2020 Machine Learning for Congestion Management and Routability Prediction within FPGA Placement
abstract
Placement for Field Programmable Gate Arrays (FPGAs) is one of the most important but time-consuming steps for achieving design closure. This article proposes the integration of three unique machine learning models into the state-of-the-art analytic placement tool GPlace3.0 with the aim of significantly reducing placement runtimes. The first model, MLCong, is based on linear regression and replaces the computationally expensive global router currently used in GPlace3.0 to estimate switch-level congestion. The second model, DLManage, is a convolutional encoder-decoder that uses heat maps based on the switch-level congestion estimates produced by MLCong to dynamically determine the amount of inflation to apply to each switch to resolve congestion. The third model, DLRoute, is a convolutional neural network that uses the previous heat maps to predict whether or not a placement solution is routable. Once a placement solution is determined to be routable, further optimization may be avoided, leading to improved runtimes. Experimental results obtained using 372 benchmarks provided by Xilinx Inc. show that when all three models are integrated into GPlace3.0, placement runtimes decrease by an average of 48%.
Hannah Szentimrey, Abeer Alhyari, Jérémy Foxcroft, Timothy Martin, David Noel, Gary William Grewal, Shawki Areibi
ACM Trans. Design Autom. Electr. Syst.4
2019 A Flat Timing-Driven Placement Flow for Modern FPGAs
abstract
In this paper, we propose a novel, flat analytic timing-driven placer without explicit packing for Xilinx UltraScale FPGA devices. Our work uses novel methods to simultaneously optimize for timing, wirelength and congestion throughout the global and detailed placement stages. We evaluate the effectiveness of the flat placer on the ISPD 2016 benchmark suite for the xcvu095 UltraScale device, as well as on industrial benchmarks. Experimental results show that on average, FTPlace achieves an 8% increase in maximum clock rate, an 18% decrease in routed wirelength, and produces placements that require 80% less time to route when compared to Xilinx Vivado 2018.1.
Timothy Martin, Dani Maarouf, Ziad Abuowaimer, Abeer Alhyari, Gary William Grewal, Shawki Areibi
DAC1
2019 Novel Congestion-estimation and Routability-prediction Methods based on Machine Learning for Modern FPGAs
abstract
Effectively estimating and managing congestion during placement can save substantial placement and routing runtime. In this article, we present a machine-learning model for accurately and efficiently estimating congestion during FPGA placement. Compared with the state-of-the-art machine-learning congestion-estimation model, our results show a 25% improvement in prediction accuracy. This makes our model competitive with congestion estimates produced using a global router. However, our model runs, on average, 291× faster than the global router. Overall, we are able to reduce placement runtimes by 17% and router runtimes by 19%. An additional machine-learning model is also presented that uses the output of the first congestion-estimation model to determine whether or not a placement is routable. This second model has an accuracy in the range of 93% to 98%, depending on the classification algorithm used to implement the learning model, and runtimes of a few milliseconds, thus making it suitable for inclusion in any placer with no worry of additional computational overhead.
Abeer Alhyari, Ziad Abuowaimer, Timothy Martin, Gary William Grewal, Shawki Areibi, Anthony Vannelli
ACM Trans. Reconfigurable Technol. Syst.3
2018 Machine-Learning Based Congestion Estimation for Modern FPGAs
abstract
Avoiding congestion for routing resources has become one of the most important placement objectives. In this paper, we present a machine-learning model for accurately and efficiently estimating congestion during FPGA placement. Compared with the state-of-the-art machine-learning congestion-estimation model, our results show a 25% improvement in prediction accuracy. This makes our model competitive with congestion estimates produced using a global router. However, our model runs, on average, 291x faster than the global router.
Dani Maarouf, Abeer Alhyari, Ziad Abuowaimer, Timothy Martin, Andrew David Gunter, Gary William Grewal, Shawki Areibi, Anthony Vannelli
FPL4
2018 GPlace3.0: Routability-Driven Analytic Placer for UltraScale FPGA Architectures
abstract
Optimizing for routability during FPGA placement is becoming increasingly important, as failure to spread and resolve congestion hotspots throughout the chip, especially in the case of large designs, may result in placements that either cannot be routed or that require the router to work excessively hard to obtain success. In this article, we introduce a new, analytic routability-aware placement algorithm for Xilinx UltraScale FPGA architectures. The proposed algorithm, called GPlace3.0, seeks to optimize both wirelength and routability. Our work contains several unique features including a novel window-based procedure for satisfying legality constraints in lieu of packing, an accurate congestion estimation method based on modifications to the pathfinder global router, and a novel detailed placement algorithm that optimizes both wirelength and external pin count. Experimental results show that compared to the top three winners at the recent ISPD’16 FPGA placement contest, GPlace3.0 is able to achieve (on average) a 7.53%, 15.15%, and 33.50% reduction in routed wirelength, respectively, while requiring less overall runtime. As well, an additional 360 benchmarks were provided directly from Xilinx Inc. These benchmarks were used to compare GPlace3.0 to the most recently improved versions of the first- and second-place contest winners. Subsequent experimental results show that GPlace3.0 is able to outperform the improved placers in a variety of areas including number of best solutions found, fewest number of benchmarks that cannot be routed, runtime required to perform placement, and runtime required to perform routing.
Ziad Abuowaimer, Dani Maarouf, Timothy Martin, Jérémy Foxcroft, Gary William Grewal, Shawki Areibi, Anthony Vannelli
ACM Trans. Design Autom. Electr. Syst.3