Dinesh Gaitonde

dblp:158/8088 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0001-8823-9689ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 PathSteiner: Improving PathFinder with Quasi-Optimal Steiner-Tree Initialization
Shashwat Shrivastava, Luka Kuresevic, Alexandros Poupakis, Chirag Ravishankar, Dinesh Gaitonde, Stefan Nikolic 0001, Mirjana Stojilovic
FCCM5
2025 Guaranteed Yet Hard to Find: Uncovering FPGA Routing Convergence Paradox
abstract
Routing is one of the major challenges of FPGA compilation. PathFinder is a ubiquitous FPGA routing algorithm used in industry and academia due to its ability to adapt to arbitrary routing architectures and user circuits. However, to this day, we do not fully understand why PathFinder works so well and what its limitations are. When a circuit fails to route, it is difficult to pinpoint the problem: architecture or algorithm. Usually, in such cases, either Pathfinder is fine-tuned or routing resources are added to the architecture to improve routability, thereby ignoring the inherent inefficiencies that may exist in Pathfinder, which further prevents us from designing silicon-efficient architectures. In this work, to pinpoint the problem, we construct constrained routing problems where nets have access to limited but specific routing resources that guarantee a legal routing solution. Yet, even with a state-of-the-art implementation, PathFinder fails to find the guaranteed routing solution or any other solution, highlighting issues specific to PathFinder. The reduced search space makes the underlying behavior more accessible for analysis and reasoning, allowing us to identify inefficiencies in the current PathFinder paradigm and propose a solution to address them. We uncovered that PathFinder's greedy approach of routing individual connections yields an inefficient route tree, in terms of the total number of nodes. We then transfer this insight-through a simple yet effective algorithm-from the constrained to the standard setting, where the search space is not reduced. By constructing more efficient route trees, the routed wirelength and the number of routed connections were reduced by 6.4% and 3.6%, respectively, on average.
Shashwat Shrivastava, Stefan Nikolic 0001, Sun Tanaka, Chirag Ravishankar, Dinesh Gaitonde, Mirjana Stojilovic
FCCM5
2023 Modular and Lean Architecture with Elasticity for Sparse Matrix Vector Multiplication on FPGAs
abstract
The use of domain-specific accelerators is becoming prominent for a variety of emerging domains such as graph analytics and HPC, where most of the computations revolve around Sparse Matrix-Vector (SpMV) Multiplication. Many of the existing SpMV accelerators do not scale well on FPGA fabric and exhibit significant performance and area overheads [1]–[3]. With the increased external memory bandwidths supported by FPGA platforms, SpMV accelerator design sizes are growing rapidly and exhibit timing closure challenges in physical implementation [4], [5]. To utilize all the High Bandwidth Memory (HBM) channels on the FPGA device, accelerator designers rely on the reuse and replication of the processing elements (PEs). As the number of PEs in a design grows, the achieved frequency of these large designs is often much lower than a single PE design [4], [5]. In this paper, we present a modular and lean architecture for the SpMV workload enabling elastic communication between building blocks. The proposed SpMV accelerator uses single-precision floating-point arithmetic (FP32) and achieves a frequency of 465 MHz for single-instance implementation. The lean nature of the design enables the scaling of the accelerator to sixteen instances, which utilizes all of the 32 HBM pseudo-channels available on the Alveo U280 FPGA platform. The accelerator design with sixteen SpMV instances, spanning multiple FPGA dies, can close timing at 310 MHz which is 80% higher than GraphLily [4] and 40% higher than HiSparse [5]. We demonstrate up to 50 GFLOPS performance on the Alveo U280 FPGA Platform which is 2.5× of GraphLily [4].
Abhishek Kumar Jain, Chirag Ravishankar, Hossein Omidian, Sharan Kumar, Maithilee Kulkarni, Aashish Tripathi, Dinesh Gaitonde
FCCM7
2023 Mitigating the Last-Mile Bottleneck: A Two-Step Approach For Faster Commercial FPGA Routing
abstract
We identified that in modern commercial FPGAs, routing signals from the general interconnect to the inputs of the CLB primitives through a very sparse input interconnect block (IIB) represents a significant runtime bottleneck. This is despite academic research often neglecting the runtime of last-mile routing through the IIB. We propose a two-step routing approach that allows resolving this bottleneck by leveraging massive parallelism of today's compute infrastructure. The main premise that enables massive parallelization is that once the signals are legally routed in the general interconnect-only reaching the inputs of the IIB, but not the final targets-the remaining last-mile routing through the IIB can be completed independently for each FPGA tile.
Shashwat Shrivastava, Stefan Nikolic 0001, Chirag Ravishankar, Dinesh Gaitonde, Mirjana Stojilovic
FPGA4
2023 AMD Next-Generation FPGA Built from Chiplets
Dinesh Gaitonde
HCS1
2023 IIBLAST: Speeding Up Commercial FPGA Routing by Decoupling and Mitigating the Intra-CLB Bottleneck
abstract
We identified that in modern commercial FPGAs, routing signals from the general interconnect to the configurable logic blocks (CLBs) through a very sparse input interconnect block (IIB) represents a significant runtime bottleneck. This is despite academic research usually altogether neglecting the runtime of last-mile routing through the IIB. To alleviate this bottleneck, we combine computer-aided design (CAD) and FPGA architecture enhancements. We propose a multi-stage FPGA routing approach, based on the premise that once the signals are legally routed in general interconnect-only reaching the inputs of the IIB, but not the final targets-the remaining last-mile routing through the IIB can be completed efficiently and independently for each FPGA tile. Then, the final routing solution can simply be built by joining the previously obtained partial solutions. However, we observe that some properties of modern IIB architectures limit the success rate of the intra-CLB routing, creating the need for revisiting the routing in the general interconnect and inevitably impairing the multi-stage routing runtime gains. We show that an enhanced IIB architecture mitigates the issue at a minimal cost. With ISPD16 benchmarks and an FPGA architecture model closely resembling AMD UltraScale FPGAs, we demonstrate the dominant contribution of last-mile routing to the router's runtime. After applying our multi-stage routing approach and the proposed enhancements, we show that the observed bottleneck can be mitigated, resulting in 4.94× faster routing on average.
Shashwat Shrivastava, Stefan Nikolic 0001, Chirag Ravishankar, Dinesh Gaitonde, Mirjana Stojilovic
ICCAD4
2020 A Domain-Specific Architecture for Accelerating Sparse Matrix Vector Multiplication on FPGAs
abstract
FPGAs allow custom memory hierarchy and flexible data movement with highly fine-grained control. These capabilities are critical for building high performance and energy efficient domain-specific architectures (DSAs), especially for workloads with irregular memory access and data-dependent communication patterns. Sparse linear algebra operations, especially sparse matrix vector multiplication (SpMV), are examples of such workloads and are becoming important due to their use in numerous areas of science and engineering. Existing FPGA-based DSAs for SpMV do not allow customization through plug and play of the building blocks. For example, most of these DSAs require switching network/crossbar architecture as a building block for routing matrix data to banked vector memory blocks. In this paper, we first present an approach where a custom network is built using simple blocks arranged in a regular fashion to exploit low-level architecture details. Further, we make use of this network to replace expensive crossbars employed in GEMX SpMV engine and develop an end-to-end tool-flow around mixed IP approach (HLS/RTL). Due to the modularity of our design, our tool-flow allows us to insert an additional block in the design to guarantee zero-stall from the accumulation stage. On Alveo U200, we report performance numbers of up to 4.4 GFLOPS (92% peak bandwidth utilization) using our accelerator (attached with one DDR4).
Abhishek Kumar Jain, Hossein Omidian, Henri Fraisse, Mansimran Benipal, Lisa Liu, Dinesh Gaitonde
FPL6
2019 Xilinx Adaptive Compute Acceleration Platform: VersalTM Architecture
abstract
In this paper we describe Xilinx's Versal-Adaptive Compute Acceleration Platform (ACAP). ACAP is a hybrid compute platform that tightly integrates traditional FPGA programmable fabric, software programmable processors and software programmable accelerator engines. ACAP improves over the programmability of traditional reconfigurable platforms by introducing newer compute models in the form of software programmable accelerators and by separating out the data movement architecture from the compute architecture. The Versal architecture includes a host of new capabilities, including a chip-pervasive programmable Network-on-Chip (NoC), Imux Registers, compute shell, more advanced SSIT, adaptive deskew of global clocks, faster configuration, and other new programmable elements as well as enhancements to the CLB and interconnect. We discuss these architectural developments and highlight their key motivations and differences in relation to traditional FPGA architectures.
Brian Gaide, Dinesh Gaitonde, Chirag Ravishankar, Trevor Bauer
FPGA2
2019 Network-on-Chip Programmable Platform in VersalTM ACAP Architecture
abstract
This paper outlines the Network-on-Chip (NoC) on Xilinx's next generation Versal-architecture. It is a hardened NoC that is present in Xilinx's next-generation 7nm architecture devices. These devices include many other new hardened features that make up the Adaptable Computing Acceleration Platform (ACAP) devices. There is a trend in FPGA devices of hardening many commonly used components such as processors, memory controllers and other IO controllers. The next generation of Xilinx devices take this a step further by providing a device-global memory mapped NoC which connects these components and the fabric in an integrated fashion. The NoC unifies communication between the processor system, FPGA fabric, memory subsystem and other hardened accelerator functions. This paper gives an overview of the Versal architecture NoC. It also motivates some of the specific characteristics of the architecture. We show how hardening the NoC lets users quickly implement high performance system level interconnect.
Ian Swarbrick, Dinesh Gaitonde, Sagheer Ahmad, Brian Gaide, Ygal Arbel
FPGA2
2018 A SAT-based Timing Driven Place and Route Flow for Critical Soft IP
abstract
Many FPGA designs contain soft IP tightly connected to hard blocks such as on-chip Processor, PCIE or IOs. Generally, these soft IPs pose significant timing closure challenges. In this paper, we propose a timing-driven Place and Route flow based on Boolean Satisfiability (SAT). Its main advantages over previous SAT-based approaches are its improved scalability and its timing awareness. We validate our flow using an IP targeting the emulation market. We demonstrate that our flow can significantly improve the usable bandwidth of FPGA IOs. Since the proposed flow is SAT based, the performance does not depend on specific ways in which more traditional place and route are usually tuned.
Henri Fraisse, Dinesh Gaitonde
FPL2
2018 Placement Strategies for 2.5D FPGA Fabric Architectures
abstract
FPGAs take advantage of 2.5D stacking technology to manufacture large capacity and high performance heterogeneous devices at reasonable costs. EDA tools need to be aware of and exploit physical characteristics of such devices, for example the reduced connection count between SLRs, the infrequency of SLL channel occurence in the fabric, and the aspect ratios of individual SLRs. We implement a partition driven placer to explore various EDA options to take advantage of architectural features in 2.5D FPGAs. We improve the routability of designs by optimizing the placer for discrete SLL channels and reduced connection counts. We propose a cut schedule for the partitioner to orient the placement with awareness of the aspect ratio of SLRs to improve track demands within each SLR.
Chirag Ravishankar, Dinesh Gaitonde, Trevor Bauer
FPL2
2018 SAT Based Place-And-Route for High-Speed Designs on 2.5D FPGAs
abstract
2.5D stacking technology allows us to build high performance and high capacity FPGA devices at reasonable costs. The communication between multiple dies happen on a passive silicon interposer at high speed, which pose several interesting challenges. Due to clock skew characteristics across multiple dies and increase in the min-max spread of delays, place-and-route tools need to address inter-die hold violations and optimize for performance. We implement a tractable SAT based methodology to achieve this by minimally detouring data paths to meet all hold requirements while optimizing performance. We also confine the solution to a small window around each inter-die (Laguna) channel to reduce routing resource utilization, congestion, and scale the methodology to any Laguna channel utilization. We improve performance across the interface by 11% compared to a state-of-the-art commercial flow and meet a 500MHz spec on Xilinx(R) UltraScale+TMdevices in 2E speedgrade. We address the scalability concerns of SAT and show how we can use this in practice with negligible runtimes in implementation tools. Our solution paves the way for FPGA-as-a-service platforms where fast inter-die communication, that does not interfere with user specific logic, is pivotal to their success.
Chirag Ravishankar, Henri Fraisse, Dinesh Gaitonde
FPT3
2016 Boolean Satisfiability-Based Routing and Its Application to Xilinx UltraScale Clock Network
abstract
Boolean Satisfiability (SAT)-based routing offers a unique advantage over conventional routing algorithms by providing an exhaustive approach to find a solution. Despite that advantage, commercial FPGA CAD tools rarely use SAT-based routers due to scalability issues. In this paper, we revisit SAT-based routing and propose two SAT formulations independent of routing architecture. We then demonstrate that SAT-based routing using either formulation dramatically outperforms conventional routing algorithms in both runtime and robustness for the clock routing of Xilinx UltraScale devices. Finally, we experimentally show that one of the proposed SAT formulations leads to a routing 18x faster and produces formulas 20x more compact than the other. This framework has been implemented into Vivado and is now currently used in production.
Henri Fraisse, Dinesh Gaitonde, Alireza Kaviani
FPGA3
2015 Enhancements in UltraScale CLB Architecture
abstract
Each generation of FPGA architecture benefits from optimizations around its technology node and target usage. In this paper, we discuss some of the changes made to the CLB for Xilinx's 20nm UltraScale product family. We motivate those changes and demonstrate better results than previous CLB architectures on a variety of metrics. We show that, in demanding scenarios, logic placed in an UltraScale device requires 16% less wirelength than 7-series. Designs mapped to UltraScale devices also require fewer logic tiles. In this paper, we demonstrate the utilization benefits of the UltraScale CLB attributed to certain CLB enhancements. The enhancements described herein result in an average packing improvement of 3% for the example design suite. We also show that the UltraScale architecture handles aggressive, tighter packing more gracefully than previous generations of FPGA. These significant reductions in wirelength and CLB counts translate directly into power, performance and ease-of-use benefits.
Shant Chandrakar, Dinesh Gaitonde, Trevor Bauer
FPGA2
2014 High capacity and high performance 20nm FPGAs
abstract
This article consists of a collection of slides from the author's conference presentation on the special features, system design and architectures, processing capabilities, and targeted markets for Xilinx's UltraSCALE, a high capacity and high performance 20 nm FPGA family of products.
Dinesh Gaitonde
Hot Chips Symposium2