Amin Mohaghegh

dblp:252/5981 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0006-7953-8963ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Déjà Vu Packing: Optimizing FPGA Logic Clustering Runtime via Pattern Memoization
abstract
Implementing a digital circuit on a field-programmable gate array (FPGA) fabric requires clustering technology-mapped netlist primitives into coarser-granularity blocks that can be directly mapped to the physical resources available on the FPGA fabric. As the internal architecture of FPGA logic blocks (LBs) has grown in complexity, with sophisticated logic elements (LEs) and highly irregular local interconnect, this packing problem has become significantly more challenging. To ensure the feasibility of intracluster routing, the computer-aided design (CAD) tools must solve a costly multi-source multi-sink routing problem for each candidate cluster. In this paper, we first show that such packing legality checks consume a significant portion of the CAD flow runtime for LB architectures with complex LEs and local routing structures resembling modern commercial FPGAs. We demonstrate that the packing stage constitutes 58% and 94% of the entire Versatile Place and Route (VPR) flow runtime on average when mapping a wide variety of benchmarks to the AMD 7-series-like and Altera Stratix 10–like VTR architecture captures, respectively. By analyzing the packing algorithm behavior, we observe that a significant fraction of the attempted packed clusters are repetitions of a much smaller number of packing patterns, and therefore many of the packing legality checks are redundant and could be skipped. To this end, we introduce our Déjà Vu packing approach, which leverages a novel packing signature tree data structure that enables efficient identification of recurring packing patterns and memoization of their legality check outcomes. Our approach speeds up the packing runtime by up to 13.4× and 29.3×, with an average of 3.7× and 6.9×, across the evaluated benchmarks on the 7-series and Stratix 10 architecture captures. These packing runtime gains result in a significant 1.6× and 5.3× average reduction in end-to-end VPR runtime, while maintaining quality of results.
Milo Liebster, Amin Mohaghegh, Andrew Boutros
FCCM2
2025 VTR 9: Open-Source CAD for Fabric and Beyond FPGA Architecture Exploration
abstract
This work details the capabilities of a major new release of the Verilog-to-Routing (VTR) open source FPGA CAD tool flow. Enhancements include generalizations of VTR’s architecture modeling language and optimizers to enable a more diverse set of programmable routing fabrics, FPGAs with embedded hard Networks-on-Chip (NoCs) and three-dimensional 3D FPGA systems that leverage stacked silicon integration. The new Parmys logic synthesis flow improves language coverage and result quality, and the physical implementation flow includes a more efficient placement engine, floorplanning constraints to guide placement, the ability to perform single-stage (flat) routing to improve quality, and parallel routing algorithms to reduce CPU time. This release also includes new architecture captures of recent commercial devices (Xilinx’s 7-series and Altera’s Stratix 10) and new benchmark suites (Titanium25 and Hermes) to aid FPGA architecture investigation. Verilog language coverage is greatly improved with the new Parmys logic synthesis flow, enabling more designs to be used with VTR. Finally, the placement and routing engines have beeenbeen sped up by 4 \(\times\) and 2.2 \(\times\) vs. VTR 8, respectively, leading to an overall physical implementation flow CPU time reduction of 48% with better result quality on average compared to VTR 8.
Mohamed A. Elgammal, Amin Mohaghegh, Soheil Gholami Shahrouz, Fatemehsadat Mahmoudi, Fahrican Kosar, Kimia Talaei, Joshua Fife, Daniel Khadivi, Kevin E. Murray, Andrew Boutros, Kenneth B. Kent, Jeffrey B. Goeders, Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.2
2025 Corrigendum: VTR 9: Open-Source CAD for Fabric and Beyond FPGA Architecture Exploration
abstract
This is a corrigendum for the article “VTR 9: Open-Source CAD for Fabric and Beyond FPGA Architecture Exploration” published in ACM Trans. Reconfig. Technol. Syst. 18, 3, Article 39 (August 2025), 53 pages.
Mohamed A. Elgammal, Amin Mohaghegh, Soheil Gholami Shahrouz, Fatemehsadat Mahmoudi, Fahrican Kosar, Kimia Talaei, Joshua Fife, Daniel Khadivi, Kevin E. Murray, Andrew Boutros, Kenneth B. Kent, Jeffrey B. Goeders, Vaughn Betz
ACM Trans. Reconfigurable Technol. Syst.2
2023 Tear Down The Wall: Unified and Efficient Intra-and Inter-Cluster Routing for FPGAs
abstract
Routing is one of the most time-consuming phases of the FPGA CAD flow, and its quality of result greatly impacts design speed and whether a successful implementation is achieved with the fixed wiring resources available. Routing is also strongly affected by the FPGA programmable interconnect architecture, making data-driven algorithms crucial to explore new fabrics. The VTR CAD flow has historically split the routing problem into two parts: within the logic cluster and between clusters. This allows a flexible specification of the programmable interconnect of each of these components and also simplifies the routing problem by splitting it into two sub-problems. However, for some architectures, this split can significantly reduce result quality, so in this work, we develop a single-stage router, run-flat, that can simultaneously optimize the intra-cluster and inter-cluster routing. We show that this new router increases the generality of architectures VTR can target and improves the quality of results at the cost of a run time increase. We detail router enhancements that reduce memory by removing non-essential nodes from the routing architecture, reduce run time by accurately ranking partial routing choices with an enhanced lookahead and efficiently resolve resource overuse at choke points with a new net-aware congestion model. Over the VTR benchmarks on an architecture with partial crossbars within the logic clusters, run-flat reduces the minimum channel width by 14% and critical path delay by 2% on average at the cost of 2.7 xthe router run time. On the Titan benchmark suite on a Stratix-IV-like architecture, run-flat reduces wirelength by 6% while taking 1.36 x the router run time.
Amin Mohaghegh, Vaughn Betz
FPL1