VLDB 2026 Research / reviewers in the wild / expert
Mohamed A. Elgammal
dblp:235/0082
· DBLP profile ↗
7ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0001-8555-7331ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Errors to Solutions: LLM-Powered Command Scripting for FPGA Cad ToolsabstractComputer-aided design (CAD) tools provide hundreds or even thousands of options that control various optimizations throughout the design flow. While this flexibility is powerful, it requires significant experience to be familiar with those options and effectively utilize them. For example, when a design fails, in many cases errors can be resolved by adjusting the CAD tool options rather than modifying the design itself. In this work, we propose VPR-LLM, a tool that utilizes Large Language Models (LLMs) to automate error resolution in the open-source FPGA CAD tool Verilog-to-Routing (VTR) by modifying the command-line options used to run the tool. VPRLLM parses error logs, VTR help messages, and documentation, then utilizes an LLM to generate modified command-line options that resolve the issue. VPR-LLM supports various LLM models and prompting techniques. All these models and techniques are evaluated and compared in terms of efficiency and cost. To evaluate our method, we proposed a dataset of 26 VTR run failures spanning five distinct error categories. The proposed technique successfully resolved 80 % of the cases without requiring any fine-tuning to the LLM model, demonstrating the effectiveness of VPR-LLM. This work represents an initial step toward AI-assisted debugging in CAD flows, where LLMs can enhance productivity by automatically identifying and correcting tool configurations. Mohamed A. Elgammal, Vaughn Betz |
FPL | 1 |
| 2025 | VTR 9: Open-Source CAD for Fabric and Beyond FPGA Architecture ExplorationabstractThis work details the capabilities of a major new release of the Verilog-to-Routing (VTR) open source FPGA CAD tool flow. Enhancements include generalizations of VTR’s architecture modeling language and optimizers to enable a more diverse set of programmable routing fabrics, FPGAs with embedded hard Networks-on-Chip (NoCs) and three-dimensional 3D FPGA systems that leverage stacked silicon integration. The new Parmys logic synthesis flow improves language coverage and result quality, and the physical implementation flow includes a more efficient placement engine, floorplanning constraints to guide placement, the ability to perform single-stage (flat) routing to improve quality, and parallel routing algorithms to reduce CPU time. This release also includes new architecture captures of recent commercial devices (Xilinx’s 7-series and Altera’s Stratix 10) and new benchmark suites (Titanium25 and Hermes) to aid FPGA architecture investigation. Verilog language coverage is greatly improved with the new Parmys logic synthesis flow, enabling more designs to be used with VTR. Finally, the placement and routing engines have beeenbeen sped up by 4 \(\times\) and 2.2 \(\times\) vs. VTR 8, respectively, leading to an overall physical implementation flow CPU time reduction of 48% with better result quality on average compared to VTR 8. Mohamed A. Elgammal, Amin Mohaghegh, Soheil Gholami Shahrouz, Fatemehsadat Mahmoudi, Fahrican Kosar, Kimia Talaei, Joshua Fife, Daniel Khadivi, Kevin E. Murray, Andrew Boutros, Kenneth B. Kent, Jeffrey B. Goeders, Vaughn Betz |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2025 | Corrigendum: VTR 9: Open-Source CAD for Fabric and Beyond FPGA Architecture ExplorationabstractThis is a corrigendum for the article “VTR 9: Open-Source CAD for Fabric and Beyond FPGA Architecture Exploration” published in ACM Trans. Reconfig. Technol. Syst. 18, 3, Article 39 (August 2025), 53 pages. Mohamed A. Elgammal, Amin Mohaghegh, Soheil Gholami Shahrouz, Fatemehsadat Mahmoudi, Fahrican Kosar, Kimia Talaei, Joshua Fife, Daniel Khadivi, Kevin E. Murray, Andrew Boutros, Kenneth B. Kent, Jeffrey B. Goeders, Vaughn Betz |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2023 | VPR-Gym: A Platform for Exploring AI Techniques in FPGA Placement OptimizationabstractWith the increasing complexity and capacity of modern Field-Programmable Gate Arrays (FPGAs), there is a growing demand for efficient FPGA computer-aided design (CAD) tools, particularly in the placement stage. While some previous works, such as RLPlace, have explored the efficacy of single-state Reinforcement Learning (RL) to optimize FPGA placement by framing it as a multi-armed bandit (MAB) problem, numerous AI techniques remain unexplored due to the outstanding engineering challenges to integrate them into the FPGA CAD flow which is based on C++. In this paper, we present VPR-Gym, a Python environment built on OpenAI Gym, that allows seamless integration with various machine learning libraries including PyTorch, TensorFlow, and Nevergrad while enabling the comparison between different AI techniques for FPGA placement. Moreover, we introduce a learning objective that reformulates the FPGA placement task as an optimization problem, thereby expanding the range of AI techniques that can be investigated beyond those for MAB problems. To showcase the capabilities of our platform, we conduct experiments comparing the performance of various MAB algorithms and evolution strategy (ES) algorithms. Our findings demonstrate that the ES approaches exhibit superior performance over the existing MAB approaches, highlighting the effectiveness of VPR-Gym in facilitating AI research to enhance FPGA placement. Ruichen Chen, Shengyao Lu, Mohamed A. Elgammal, Peter Chun, Vaughn Betz, Di Niu 0002 |
FPL | 3 |
| 2023 | Koios 2.0: Open-Source Deep Learning Benchmarks for FPGA Architecture and CAD Researchabstractthe prevalence of deep learning (DL) in many applications, researchers are investigating different ways of optimizing field-programmable gate array (FPGA) architecture and CAD to achieve better quality-of-results (QoRs) on DL-based workloads. In this optimization process, benchmark circuits are an essential component; the QoR achieved on a set of benchmarks is the main driver for architecture and CAD design choices. However, current academic benchmark suites are inadequate, as they do not capture any designs from the DL domain. This work presents the second version of our suite of DL acceleration benchmark circuits for FPGA architecture and CAD research, called Koios. This suite of 40 circuits covers a wide variety of accelerated neural networks, design sizes, implementation styles, abstraction levels, and numerical precisions. These benchmarks include 32 DL designs and eight synthetic (proxy) benchmarks. The Koios benchmarks are larger, more data parallel, more heterogeneous, more deeply pipelined, and utilize more FPGA architectural features compared to existing open-source benchmarks. This enables researchers to pinpoint architectural inefficiencies for this class of workloads and optimize CAD tools on more representative benchmarks that stress the CAD algorithms in different ways. In this article, we describe the Koios designs, compare their characteristics to prior FPGA benchmark suites, and present results of running them through the verilog-to-routing (VTR) flow using a recent FPGA architecture model. Finally, we present case studies showing how exploration of DL-optimized FPGA architecture and CAD algorithms can be performed using our new benchmark suite. Aman Arora 0001, Andrew Boutros, Seyed Alireza Damghani, Karan Mathur, Vedant Mohanty, Tanmay Anand, Mohamed A. Elgammal, Kenneth B. Kent, Vaughn Betz, Lizy Kurian John |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2022 | Quality & Generality: A Flexible FPGA Re-Clustering Technique to Improve Packing and PlacementabstractThe Packing and Placement stages are two major steps in the FPGA backend flow which greatly affect the Quality-of-Results (QoR) of design implementation. While these problems have been extensively studied in the literature, most approaches have either sacrificed generality by targeting specific and simplified FPGAs with few “block packing” legality constraints, or sacrificed quality by making irreversible packing decisions early in the flow and hence constraining the optimizations available to the subsequent placement stage. In this paper, we propose a new (re-clustering API) that can be used to update the packed netlist at different points throughout the packing and placement stages. This API can be used in our proposed flow to improve the QoR while preserving the generality and flexibility of the flow and ensuring the legality of the solution for any proposed FPGA architecture. Mohamed A. Elgammal, Vaughn Betz |
FPT | 1 |
| 2022 | RLPlace: Using Reinforcement Learning and Smart Perturbations to Optimize FPGA PlacementabstractSimulated annealing (SA) is one of the most common FPGA placement techniques, and is used both as a standalone algorithm and to improve an initial analytical placement. While SA-based placers can achieve high-quality results, they suffer from long runtimes. In this article, we introduceRLPlace, a novel SA-based FPGA placer that utilizes both reinforcement learning (RL) and targeted perturbations (directed moves). The proposed moves target both wirelength and timing optimization and explore the solution space more efficiently than traditional random moves while preventing oscillation in the Quality of Results (QoR). RL techniques are used to dynamically select the most effective move types as optimization progresses. The experimental results show thatRLPlaceoutperforms the widely used VTR 8 placer across all runtime/quality tradeoff points, achieving better QoR placement solutions in less runtime. On average, across the Titan23 suite of large FPGA benchmarks, RLPlace can reduce CPU time by$2.5\times $with result quality comparable to VTR 8, or improve wirelength by 8% (at a high CPU time budget) −26% (at a low CPU time budget) versus VTR 8.0 given the same CPU time. Mohamed A. Elgammal, Kevin E. Murray, Vaughn Betz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |