VLDB 2026 Research / reviewers in the wild / expert
Grace Zgheib
dblp:33/9240
· DBLP profile ↗
23ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-1476-2984ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 5 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating Multi-agent Reinforcement Learning on Heterogeneous Platforms
Samuel Wiggins, Grace Zgheib, Mahesh A. Iyer, Viktor Prasanna 0001 |
Euro-Par (1) | 2 |
| 2026 | FAME: A Framework for Accelerating Independent Multi-Agent Reinforcement Learning on Heterogeneous PlatformsabstractMulti-Agent Reinforcement Learning (MARL) enables multiple autonomous agents to learn and act in a shared environment. Independent learning (IL) is a widely used MARL paradigm that underpins many real-world applications requiring efficient training at scale. However, accelerating IL at scale is non-trivial. Existing MARL frameworks rely on single-process execution and homogeneous hardware assumptions, limiting scalability and underutilizing modern heterogeneous platforms composed of CPUs, GPUs, and FPGAs. Addressing this gap requires new execution models that increase parallelism while preserving IL training semantics. In this work, we present FAME, a framework that distributes computation across heterogeneous hardware resources while providing flexible interfaces that allow MARL practitioners to prototype and test new IL approaches. FAME is composed of: (1) high-level APIs that simplify IL algorithm development, (2) a heterogeneous IL training protocol that supports concurrent agent training on multiple diverse devices, while maintaining algorithm-agnostic training semantics, (3) automatic hardware configuration generation that optimizes system throughput without needing users to manually fine-tune their system setup, and (4) dynamic load balancing among devices with different compute and memory characteristics. We demonstrate FAME’s capabilities using three representative IL algorithms on a heterogeneous node platform consisting of CPUs, GPUs, and FPGAs. Implementations generated using FAME achieve a geometric mean end-to-end training time speedup of 7.1 × over state-of-the-art implementations and up to 2.7 × speedup over additional highly parallel baselines developed in this work. Samuel Wiggins, Nikunj Gupta, Grace Zgheib, Mahesh A. Iyer, Viktor Prasanna 0001 |
HPDC | 3 |
| 2025 | Accelerating Independent Multi-Agent Reinforcement Learning on Multi-GPU Platforms
Samuel Wiggins, Nikunj Gupta, Grace Zgheib, Mahesh A. Iyer, Viktor Prasanna 0001 |
Euro-Par (3) | 3 |
| 2025 | A Partitioning-Based CAD Flow for Interposer-Based Multi-Die FPGAsabstractMulti-die interposer-based FPGA architectures present several design challenges: (i) limited inter-die connectivity (interposer) resources and (ii) increased interposer delays compared to intra-die routing. Addressing these challenges is critical, as compared to intra-die routing, they directly impact the routability, routed wirelength (rWL) and maximum clock frequency (Fmax) of the design. In this paper, we present a partitioning-based CAD flow tailored for interposer-based multi-die FPGA architectures. Central to our approach is FPGAPart, the first open-source timing-driven netlist partitioner that can handle FPGA designs while addressing modern architectural constraints. We integrate FPGAPart with the open-source tool VTR 7.0. In particular, we use: (i) VTR 7.0's pre-packing solutions as clustering hints during partitioning and (ii) the FP-Growth algorithm [14] to detect frequently occurring patterns (instances) across multiple timing paths for clustering. Additionally, we introduce neighborhood influences-based cutting planes into the core ILP solver in FPGAPart, resulting in a ~38× ILP runtime speedup with < 1 % degradation in solution quality, compared to using no neighborhood influences. Compared to the default VTR 7.0, our flow achieves a geometric mean improvement of ~3% in rWL and ~3% in Fmax for a two-die configuration, with similar improvements across other configurations. Compared to hMETIS [17], METIS [18] and TritonPart [5], FPGAPart achieves improvements up to ~4% in rWL and ~7 % in Fmax for a two-die configuration, with similar improvements across other configurations. Mahesh A. Iyer, Andrew B. Kahng, Jason Luu, Bodhisatta Pramanik, Kristofer Vorwerk, Grace Zgheib |
FCCM | 6 |
| 2025 | Double Duty: FPGA Architecture to Enable Concurrent LUT and Adder Chain UsageabstractFlexibility and customization are key strengths of Field-Programmable Gate Arrays (FPGAs) when compared to other computing devices. For instance, FPGAs can efficiently implement arbitrary-precision arithmetic operations, and can perform aggressive synthesis optimizations to eliminate ineffectual operations. Motivated by sparsity and mixed-precision in deep neural networks (DNNs), we investigate how to optimize the current logic block architecture to increase its arithmetic density. We find that modern FPGA logic block architectures prevent the independent use of adder chains, and instead only allow adder chain inputs to be fed by look-up table (LUT) outputs. This only allows one of the two primitives—either adders or LUTs—to be used independently in one logic element and prevents their concurrent use, hampering area optimizations. In this work, we propose the Double Duty logic block architecture to enable the concurrent use of the adders and LUTs within a logic element. Without adding expensive logic cluster inputs, we use 4 of the existing inputs to bypass the LUTs and connect directly to the adder chain inputs. We accurately model our changes at both the circuit and CAD levels using open-source FPGA development tools. Our experimental evaluation on a Stratix-10-like architecture demonstrates area reductions of 21.6% on adder-intensive circuits from the Kratos benchmarks, and 9.3% and 8.2% on the more general Koios and VTR benchmarks respectively. These area improvements come without an impact to critical path delay, demonstrating that higher density is feasible on modern FPGA architectures by adding more flexibility in how the adder chain is used. Averaged across all circuits from our three evaluated benchmark set, our Double Duty FPGA architecture improves area-delay product by 9.7%. Junius Pun, Xilai Dai, Grace Zgheib, Mahesh A. Iyer, Andrew Boutros, Vaughn Betz, Mohamed S. Abdelfattah |
FPL | 3 |
| 2025 | ARC: A Runtime Engine for Accelerating Independent Multi-Agent Reinforcement Learning on Multi-Core ProcessorsabstractMulti-Agent Reinforcement Learning (MARL) enables multiple agents to optimize individual or joint objectives in a shared environment, with applications spanning robotics, autonomous driving, and financial systems. Independent Learning (IL), a simple yet effective MARL approach, trains agents independently without modeling inter-agent communication or explicit coordination. This simplicity reduces computational requirements, making CPU platforms an attractive alternative to accelerators such as GPUs for smaller model architectures typical of IL. However, existing CPU-based MARL implementations rely on a Single-Learner training scheme, which sequentially trains agent networks and fails to utilize the full potential of multicore CPUs. This limits scalability and introduces inefficiencies, particularly for large-scale MARL systems. In this work, we present ARC, a lightweight runtime engine designed to accelerate IL training on multi-core CPU platforms. ARC introduces an Independent Multi-Learner training scheme that parallelizes agent model updates, maximizing hardware utilization and scalability, while preserving training semantics. By exploring and selecting optimal parallelization strategies tailored to the user's hardware, ARC ensures seamless acceleration without manual configuration. Through experiments on state-of-the-art IL algorithms, we demonstrate an increased end-to-end speedup of up to$28.2 \times$while exploring only 5% of the configuration space. We open-source ARC, supporting multiple algorithms and providing significant performance improvements, thereby facilitating the development of scalable MARL applications. Samuel Wiggins, Nikunj Gupta, Grace Zgheib, Mahesh A. Iyer, Viktor Prasanna 0001 |
ICPADS | 3 |
| 2022 | Detailed Placement for Dedicated LUT-Level FPGA InterconnectabstractIn this work, we develop timing-driven CAD support for FPGA architectures with direct connections between LUTs. We do so by proposing an efficient ILP-based detailed placer, which moves a carefully selected subset of LUTs from their original positions, so that connections of the user circuit can be appropriately aligned with the direct connections of the FPGA, reducing the circuit’s critical path delay. We discuss various aspects of making such an approach practicable, from efficient formulation of the integer programs themselves, to appropriate selection of the movable nodes. These careful considerations enable simultaneous movement of tens of LUTs with tens of candidate positions each, in a matter of minutes. In this manner, the impact of additional connections on the critical path delay more than doubles, compared to the previously reported results that relied solely on architecture-oblivious placement. Stefan Nikolic 0001, Grace Zgheib, Paolo Ienne |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2021 | Clock Skew Scheduling: Avoiding the Runtime Cost of Mixed-Integer Linear ProgrammingabstractClock Skew Scheduling has become a common practice in state-of-the-art FPGAs with the introduction of delay chains on the clock path in the hardware of both Xilinx and Intel®FPGAs, as well as clock skew scheduling algorithms in the CAD tools. Ideally, globally optimal solutions are sought to find the best solution across the entire design. However, using Mixed-Integer Linear Programming (MILP) to find such optimal solutions has a large and sometimes unrealistic runtime cost, especially for bigger designs. Besides, a high runtime does not necessarily correlate to improved performance. We present, in this paper, techniques to reduce the runtime of mixed-integer linear programming approaches to clock skew scheduling, and provide alternatives that can achieve optimal or near-optimal performance gain with a fraction of the MILP runtime cost. Grace Zgheib, Yu Shen Lu, Ilya Ganusov |
FPL | 1 |
| 2020 | Architectural Enhancements in Intel® Agilex™ FPGAsabstractThis paper describes architectural enhancements in Intel® Agilex™ FPGAs and SoCs. Agilex devices are built on Intel's 10nm process and feature next-generation programmable fabric, tightly coupled with a quad-core ARM processor subsystem, a secure device manager, IO and memory interfaces, and multiple companion transceiver tile choices. The Agilex fabric features multiple logic block enhancements that significantly improve propagation delays and integrate more effectively with the second-generation HyperFlexAgilex™ pipelined routing architecture. Routing connections are re-designed to be point-to-point, dropping intermediate connections featured in prior FPGA generations and replacing them with a wider variety of shorter wire types. Fine-grain programmable clock skew and time-borrowing were introduced throughout the fabric to augment the slack-balancing capabilities of HyperFlex registers. DSP capabilities are also extended to natively support new INT9/BFLOAT16/FP16 formats. Together, along with process and circuit enhancements, these changes support more than 40% performance improvement over the Stratix® 10 family of FPGAs. Jeffrey Chromczak, Mark Wheeler, Charles Chiasson, Dana How, Martin Langhammer, Tim Vanderhoek, Grace Zgheib, Ilya Ganusov |
FPGA | 7 |
| 2020 | Straight to the Point: Intra- and Intercluster LUT Connections to Mitigate the Delay of Programmable RoutingabstractTechnology scaling makes metal delay ever more problematic, but routing between Look-Up Tables (LUTs) still passes through a series of transistors. It seems wise to avoid the corresponding delay whenever possible. Direct connections between LUTs, both within and across multiple clusters, can eschew the transistor delays of crossbars, connection blocks, and switch blocks. In this paper we investigate the usefulness of enhancing classical Field-Programmable Gate Array (FPGA) architectures with direct connections between LUTs. We present an efficient algorithm for searching automatically the most interesting patterns of such direct connections. Despite our methods being fairly conservative and relying on the use of unmodified standard CAD tools, we obtain a 2.77% improvement of the geometric mean critical path delay of a standard benchmark set, with improvement ranging from -0.17% to 7.3% for individual circuits. As modest as these results may seem at first glance, we believe that they position direct connections between LUTs as a promising topic for future research. Extending this work with dedicated CAD algorithms and exploiting the increased possibilities for optimal buffering, diagonal routing, and pipelining could prove direct connections important to the continuation of performance improvement into next generation FPGAs. Stefan Nikolic 0001, Grace Zgheib, Paolo Ienne |
FPGA | 2 |
| 2020 | Timing-Driven Placement for FPGA Architectures with Dedicated Routing PathsabstractThe idea of introducing dedicated, fast paths between certain FPGA elements in order to reduce delay is neither new nor particularly hard to come up with. What is less obvious, however, is how to put such paths to actual use. In this work, we propose an effective ILP-based detailed placer for FPGA architectures with direct connections between LUTs. We discuss various aspects of making such an approach practicable, from efficient formulation of the integer programs themselves, to focused application of the placer on specific portions of the circuit where it could have the greatest impact. These careful considerations allow us to simultaneously move tens of LUTs with tens of candidate positions each, in a matter of minutes. This more than doubles the advantage of additional connections on the critical path delay compared to the previously reported results that relied on architecture-oblivious placement algorithms. Stefan Nikolic 0001, Grace Zgheib, Paolo Ienne |
FPL | 2 |
| 2019 | Finding a Needle in the Haystack of Hardened Interconnect PatternsabstractCircuits naturally exhibit recurring patterns of local interconnect. Hardening those patterns when designing Field Programmable Gate Array (FPGA) clusters can both eliminate slow programmable connections from the critical path and remove the need for transistors to implement them. While we may be able to manually design such clusters, based on intuition and observations, such an endeavour will always leave us in doubt whether we have used the potential of cheap and fast hardened connections to the fullest. On the other hand, since there are ~107possible patterns of interconnect among only eight 5-input Look-Up Tables (LUTs), even without considering the possibilities for enforcing input sharing, performing an exhaustive exploration of the design space seems like a task as daunting as finding a needle in a haystack. Despite their enormous sizes, design spaces spanned by such cluster architectures are very well structured. In this paper, we leverage that structure to limit the search to only those points of the space that may perform better than any chosen reference. We demonstrate the usefulness of our techniques by narrowing the space spanned by five 5-LUT structures from hundreds of millions to a set of 261 structures that achieve ? 80% utilization on a set of standard benchmarks, while requiring only 12 external inputs. We believe that the exploration techniques presented here are an important step towards truly exploiting the potentials of hardening programmable connections. Stefan Nikolic 0001, Grace Zgheib, Paolo Ienne |
FPL | 2 |
| 2017 | NAND-NOR: A Compact, Fast, and Delay Balanced FPGA Logic Element
Grace Zgheib, Wei Li 0008, Zhenghong Jiang, Kaihui Tu, Paolo Ienne, Haigang Yang |
FPGA | 3 |
| 2017 | Evaluating FPGA clusters under wide ranges of design parametersabstractThe latest published studies with extensive explorations of look-up table and cluster sizes are now more than a decade old. However, CMOS technology as well as CAD and transistor modeling tools have improved so much since that it is reasonable to wonder whether the conclusions of such studies still hold. One of the major difficulties of conducting these studies, especially in academia, is producing credible delay and area models. In this paper, we take advantage of a recently developed architecture modeling tool to re-evaluate the effect of the various cluster parameters on the FPGA. We considerably extend the exploration space beyond that of the classic studies to include sparse crossbars and fracturable LUTs, and show some results that go against the current tenets of FPGA architecture. Grace Zgheib, Paolo Ienne |
FPL | 1 |
| 2017 | Improving Circuit Mapping Performance Through MIG-based Synthesis for Carry ChainsabstractHard-wired carry chains in FPGAs are designed to improve efficiency of important arithmetic primitives. Although they are proven to be effective for arithmetic-rich functions, there are very few studies on the optimization opportunities of carry chains for general logic that is poor in arithmetic operations. Recently, Majority-Inverter Graphs (MIGs) were proposed for efficient Boolean logic optimization. MIGs open an opportunity for efficient mapping of critical paths onto hard carry chains, as the carry logic of a full adder is naturally a majority (MAJ) gate. In this paper, we propose an MIG-based synthesis method to exploit hard adders in FPGAs for the mapping of general logic. The proposed heuristic algorithm selects MAJ nodes to be mapped on the carry chains and the associated LUTs; then, the efficiency of carry chain mapping is examined theoretically for efficient LUT utilization. The experimental results show that, compared to traditional design flow Verilog-to-Routing (VTR 7.0), the proposed approach can improve delay by up to 25% with an average of 8%, while the channel width is reduced by up to 20% with an average of 6%. Zhufei Chu, Xifan Tang, Mathias Soeken, Ana Petkovska, Grace Zgheib, Luca G. Amarù, Yinshui Xia, Paolo Ienne, Giovanni De Micheli, Pierre-Emmanuel Gaillardon |
ACM Great Lakes Symposium on VLSI | 5 |
| 2016 | FPRESSO: Enabling Express Transistor-Level Exploration of FPGA ArchitecturesabstractIn theory, tools like VTR---a retargetable toolchain mapping circuits onto easily-described hypothetical FPGA architectures---could play a key role in the development of wildly innovative FPGA architectures. In practice, however, the experiments that one can conduct with these tools are severely limited by the ability of FPGA architects to produce reliable delay and area models---these depend on transistor-level design techniques which require a different set of skills. In this paper, we introduce a novel approach, which we call Fpresso, to model the delay and area of a wide range of largely different FPGA architectures quickly and with reasonable accuracy. We take inspiration from the way a standard-cell flow performs large scale transistor-size optimization and apply the same concepts to FPGAs, only at a coarser granularity. Skilled users prepare for \fpresso locally optimized libraries of basic components with a variety of driving strengths. Then, ordinary users specify arbitrary FPGA architectures as interconnects of basic components. This is globally optimized within minutes through an ordinary logic synthesis tool which chooses the most fitting version of each cell and adds buffers wherever appropriate. The resulting delay and area characteristics can be automatically used for VTR. Our results show that \fpresso provides models that are on average within some 10-20\% of those by a state-of-the-art FPGA optimization tool and is orders of magnitude faster. Although the modelling error may appear relatively high,we show that it seldom results in misranking a set of architectures, thus indicating a reasonable modeling faithfulness. Grace Zgheib, Manana Lortkipanidze, Muhsen Owaida, David Novo, Paolo Ienne |
FPGA | 1 |
| 2016 | Automatic wire modeling to explore novel FPGA architecturesabstractCorrect modeling of the FPGA architecture is a fundamental step in the design and testing of new architectures. Such modeling must include the logic components of the FPGA as well as the wires connecting these components. In recent years, researchers have been investigating tools to perform automatic modeling of FPGAs with either a fixed structure (e.g., COFFE) or even unconstrained architecture designs (e.g., FPRESSO). And, although these tools highlight the importance of intercomponent wire modeling, FPRESSO does not model wire loads due to the difficulty in determining the length of the wires without a fixed cluster structure, while COFFE approximates the wirelength by assuming a simplified topology and a predefined placement of components. In this paper, we present an automatic and generic method to not only model the intercomponent wires but also to minimize the total wirelength, by floorplanning the cluster and finding the best placement of its components. Our algorithms were integrated into FPRESSO to improve its modeling by including wire loads in its FPGA architecture modeling. Grace Zgheib, Paolo Ienne |
FPT | 1 |
| 2015 | A technology mapper for depth-constrained FPGA logic cellsabstractIn the last decade, progress in logic synthesis has brought about new advantageous circuit representations. These representations, such as And-Inverter Graphs in the ubiquitous open-source synthesizer ABC, have inspired new designs of Field Programmable Gate Arrays (FPGAs), which, instead of using Look-Up Tables (LUTs), mimic the topology of the circuit representation in the basic logic cells. More recent examples are Majority-Inverter Graphs, another uniform representation which has triggered considerable interest in synthesis and which naturally suggests new logic cells. Yet, in this paper we observe how naïvely adapting technology mapping solutions for classic LUT-based FPGAs to these new architectures incurs severe shortcomings. The key issue is that LUTs are inherently input-constrained (the logic function they implement is irrelevant) and have generally a single output; on the other hand, logic cells made of uniform networks of some fundamental logic function (e.g., And-Invert) are constrained in terms of logic depth and multiple outputs are an integral feature. We introduce novel and effective solutions to address these differences; the result is a highly versatile mapper—thus enabling further research in these new architectures—with a significantly better performance than what is described in literature for one such architecture. Specifically, when we compare with the state of the art on one sample architecture, we obtain a significant decrease in area (on average 18% over several benchmarks) while also improving slightly the critical path (a reduction of 3%). Zhenghong Jiang, Grace Zgheib, Colin Yu Lin, David Novo, Liqun Yang, Haigang Yang, Paolo Ienne |
FPL | 2 |
| 2015 | Improved carry chain mapping for the VTR flowabstractCarry chains facilitate the implementation of adders and improve the performance of arithmetic circuits in FPGAs. The last version of the commonly used open-source Verilog-to-Routing (VTR) CAD flow now enables modelling carry chains in FPGA architectures. However, one of the shortcomings of the existing flow lies in its inability to identify arithmetic operations when described as gate-level circuits. Moreover, the VTR flow squanders most of the LUTs preceding the chain logic. This paper focuses on these two problems and proposes preprocessing the circuit before technology mapping to allow for a more efficient use of carry chains. The first proposed method maps logic on the carry chains for circuits expressed using a gate-level description. On average, it identifies about 30% more meaningful full adders than the existing tool flow operating on the RTL descriptions. Area is thus improved by up to 15% with an average of 6% for almost no delay penalty. Secondly, we increase the use of the LUTs preceding the chain logic by a factor 2 on average. This reduces delay (up to 9%) and area (up to 2%), compared to the existing VTR flow. The new approach is independent of the specific carry-chain architecture and can be generically adapted to any FPGA with built-in hardened adders. Ana Petkovska, Grace Zgheib, David Novo, Muhsen Owaida, Alan Mishchenko, Paolo Ienne |
FPT | 2 |
| 2014 | Revisiting and-inverter conesabstractAnd-Invert Cones (AICs) have been suggested as an alternative to the ubiquitous Look-Up Tables (LUTs) used in commercial FPGAs. The original article suggesting the new architecture made some untested assumptions on the circuitry needed to implement AIC architectures and did not develop completely the toolset necessary to assess comprehensively the idea. In this paper, we pick up the architecture that some of us proposed in the original AIC paper and try to implement it as thoroughly as we can afford. We build all components for the logic cluster at transistor level in a 40~nm technology as well as a LUT-based architecture inspired by Altera's Stratix~IV. We first determine that the characteristics of our LUT-based architecture are reasonably similar to those of the commercial counterpart. Then, we compare the AIC architecture to the baseline on a number of benchmarks, and we find a few difficulties that had been overlooked before. We thus explore other design possibilities around the original design point and show their detailed impact. Finally, we discuss how the very structure of current logic clusters seems not perfectly appropriate for getting the best out of AICs and conclude that, even though they are not confirmed as an immediate blessing today, AICs still offer rich research opportunities. Grace Zgheib, Liqun Yang, David Novo, Hadi Parandeh-Afshar, Haigang Yang, Paolo Ienne |
FPGA | 1 |
| 2013 | Shadow AICs: reaping the benefits of and-inverter cones with minimal architectural impact (abstract only)abstractDespite their many advantages, FPGAs are still inefficient. This inefficiency is mainly due to programmable routing networks; however, FPGA logic blocks also have their share of contribution. From the performance perspective, fewer hops in the routing network translates to a shorter critical path; and that requires large logic blocks capable of covering big portions of circuits. Recent work has shown that And-Inverter Cones (AICs) can considerably reduce the number of logic block levels compared to Look-Up Tables (LUTs). The best performance is achieved when both AICs and LUTs are used, but the AIC implementation requires radical changes in the FPGAs architecture. In this paper, we use AICs as shadow logic for LUTs in LUT-clusters, which requires minimal architectural changes while exploiting the benefits of both AICs and LUTs. The basic idea is to reuse the input crossbar of LUT-clusters for the shadow AICs while combining both LUTs and AICs in the same cluster. We also propose changes in the AIC architecture to enhance mapping on AICs. Our experimental results indicate that the new cluster architecture can reduce the average circuit delay by 12% with respect to standard FPGA clusters. However, this performance gain comes at a price of 43% area overhead in terms of number of logic clusters. Our results show that for a modest 6% increase in area, FPGA manufacturers can move towards next-generation FPGA logic elements. This transition would provide faster design options without major architectural changes. Hadi Parandeh-Afshar, Grace Zgheib, David Novo, Madhura Purnaprajna, Paolo Ienne |
FPGA | 2 |
| 2013 | Shadow And-Inverter ConesabstractDespite their many advantages, FPGAs still come with significant overheads in area, delay, and power consumption due to an extreme programmability in both the routing and logic. From the performance perspective, large logic blocks, capable of covering big portions of circuits, lead to fewer hops in the routing network, and thus, to a shorter critical path. Recent work has shown that And-Inverter Cones (AICs) can considerably reduce the number of logic block levels compared to Look-Up Tables (LUTs), in a radically altered FPGAs architecture. In this paper, we use AICs as shadow logic for LUTs, which incurs minimal architectural changes with respect to current FPGAs, while exploiting the benefits of both AICs and LUTs. We also propose changes in the AIC architecture, for a more compact technology mapping. The new architecture reduces the average circuit delay by up to 35% with respect to standard FPGAs at the expense of a 3x increase in the number of the logic clusters. Other benchmarks show more moderate area overheads, e.g., 16% delay improvement for 20% area overhead. Hadi Parandeh-Afshar, Grace Zgheib, David Novo, Madhura Purnaprajna, Paolo Ienne |
FPL | 2 |
| 2011 | Reducing the pressure on routing resources of FPGAs with generic logic chainsabstractRouting resources in modern FPGAs use 50% of the silicon real estate and are significant contributors to critical path delay and power consumption; the situation gets worse with each successive process generation, as transistors scale more effectively than wires. To cope with these challenges, FPGA architects have divided wires into local and global categories and introduced fast dedicated carry chains between adjacent logic cells, which reduce routing resource usage for certain arithmetic circuits (primarily adders and subtractors). Hadi Parandeh-Afshar, Grace Zgheib, Philip Brisk, Paolo Ienne |
FPGA | 2 |