EDBT 2026 Demo / reviewers in the wild / expert
Andy Gean Ye
dblp:48/26
· DBLP profile ↗
23ranked-venue papers
9as first author
5since 2021 · last 2025
0000-0002-2959-5736ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 9 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Measuring the Minimum Power Requirement of FPGA Architectural SpecificationsabstractBypass capacitances bridge the gap between momentary surges of power demand caused by switching transistors and the ability of power supplies to respond to these surges. Even though bypass capacitances are not explicitly specified in FPGA architectural specifications, the gate capacitances of quiescent transistors in an FPGA architecture can act as symbiotic bypass capacitances. Since any additional dedicated bypass capacitances can increase the implementation area of an FPGA, it is important to measure the ability of the quiescent bypass capacitances in maintaining stable operating voltages while the FPGA is powered by voltage sources with limited power outputs. This work measures the ability of the quiescent bypass capacitances in maintaining stable supply voltages under limited power in the context of FPGA architectural investigations. We found that, for the 12-LUT logic cluster investigated in this work, to limit VDD value droop to 5% of the nominal VDD value, the power source must be able to produce 3 times of the average power consumed by the same logic cluster powered by an ideal voltage source. When the power output is further reduced, VDD value droop increases, resulting in significantly increased delay and reduced noise margin. In addition, the degree of parallel execution also has a significant effect on VDD value. In particular, at the nominal VDD value of 0.9 volts and the maximum power output of 28.1 uW, executing all 12 LUTs in parallel results in VDD value drooping to a minimum value of 0.370 volts while executing 1 LUT at a time results in VDD value drooping to a minimum of 0.690 volts, with both cases achieving similar computing time. These results suggest that FPGA architectural evaluations should take bypass capacitances and the power limit of voltage sources into consideration in order to design efficient FPGA architectures for low power applications. Andy Gean Ye, Anas Razzaq Ghumman |
FPGA | 1 |
| 2024 | Evaluating the Impact of Using Multiple-Metal Layers on the Layout Area of Switch Blocks for Tile-Based FPGAs in FinFET 7nmabstractA new area model for estimating the layout area of switch blocks is introduced in this work. The model is based on a realistic layout strategy. As a result, it not only takes into consideration the active area that is needed to construct a switch block but also the number of metal layers available and the actual dimensions of these metals. The model assigns metal layers to the routing tracks in a way that reduces the number of vias that are needed to connect different routing tracks together while maintaining the tile-based structure of FPGAs. It also takes into account the wiring area required for buffer insertion for long wire segments. The model is evaluated based on the layouts constructed in the ASAP7 FinFET 7nm Predictive Design Kit. We found that the new model, while specific to the layout strategy that it employs, improves upon the traditional active-based area estimation models by considering the growth of the metal area independently from the growth of the active area. As a result, the new model is able to more accurately estimate the layout area by predicting when the metal area will overtake the active area as the number of routing tracks is increased. This ability allows the more accurate estimation of the true layout cost of FPGA fabrics at the early floor planning and architectural exploration stage; and this increase in accuracy can encourage a wider use of custom FPGA fabrics that target specific sets of benchmarks in future SOC designs. Furthermore, our data indicate that the conclusions drawn from several significant prior architectural studies remain to be correct under FinFET geometries and wiring area considerations despite their exclusive use of active-only area models. This correctness is due to the small channel widths, around 30–60 tracks per channel, of the architectures that these studies investigate. For architectures that approach the channel width of modern commercial FPGAs with more than 100–200 tracks per channel, our data show that wiring area models justified by detailed layout considerations are an essential addition to active area models in the correct prediction of the implementation area of FPGAs. Sajjad Rostami Sani, Andy Gean Ye |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2022 | Evaluating the impact of using multiple-metal layers on the layout area of switch blocks for tile-based FPGAs in FinFET 7nmabstractA new area model for estimating the layout area of switch blocks is introduced in this work. The model not only takes into consideration the active area that is needed to construct a switch block but also the number of metal layers available and the actual dimension of these metals. The model assigns metal layers to the routing tracks in a way that reduces the number of vias that are needed to connect different routing tracks together while maintaining the tile-based structure of FPGAs. The model is evaluated based on the layouts constructed in ASAP7 FinFET 7nm Predictive Design Kit. We found that the new model improves upon the traditional active-based area estimation models by considering the growth of metal area independently from the growth of the active area. As a result, the new model is able to more accurately estimate layout area by predicting when metal area will overtake active area as the number of routing tracks is increased. This ability allows the more accurate estimation of the true layout cost of FPGA fabrics at the early floor planning and architectural exploration stage; and this increase in accuracy can encourage a wider use of custom FPGA fabrics that target specific sets of benchmarks in future SOC designs. Sajjad Rostami Sani, Anas Razzaq Ghumman, Andy Gean Ye |
FCCM | 3 |
| 2022 | The effect of gate voltage boosting on the power efficiency of multi-context FPGAs
Anas Razzaq Ghumman, Sajjad Rostami Sani, Andy Gean Ye |
Integr. | 3 |
| 2021 | Designing efficient FPGA tiles for power-constrained ultra-low-power applications
Anas Razzaq Ghumman, Sajjad Rostami Sani, Andy Gean Ye |
Integr. | 3 |
| 2020 | Measuring the Accuracy of Layout Area Estimation Models of Tile-Based FPGAs in FinFET TechnologyabstractThis work presents the layout area of encoded and decoded multiplexers, two essential building blocks of modern FPGAs, in FinFET. Layouts with both 2 and 3 metal layers based on ASAP7 Predictive Design Kit are presented. The layout area is then compared with the prediction of two equation-based models: the VPR area model and the COFFE area model. We found that, with the original model parameters which are adjusted for planar technologies, these two equation-based models are not accurate in predicting FinFET layout area with error ranges of -2.3% to +86.5% and -32.7% to +19.3% for the VPR and COFFE models, respectively. Furthermore, when the model parameters are specifically adjusted for FinFET, the error ranges remain to be large with -19% to +31% and -26.6% to +25% for the VPR and COFFE models, respectively. These data underline the importance of verifying equation-based models against actual layout areas in FPGA architectural studies, especially when there are significant changes in the underlining process technologies. Sajjad Rostami Sani, Farheen Fatima Khan, Anas Razzaq Ghumman, Andy Gean Ye |
FPL | 4 |
| 2018 | An Evaluation on the Accuracy of the Minimum-Width Transistor Area Models in Ranking the Layout Area of FPGA ArchitecturesabstractThis work provides an evaluation on the accuracy of the minimum-width transistor area models in ranking the actual layout area of FPGA architectures. Both the original VPR area model and the new COFFE area model are compared against the actual layouts with up to three metal layers for the various FPGA building blocks. We found that both models have significant variations with respect to the accuracy of their predictions across the building blocks. In particular, the original VPR model overestimates the layout area of larger buffers, full adders, and multiplexers by as much as 38%, while they underestimate the layout area of smaller buffers and multiplexers by as much as 58%, for an overall prediction error variation of 96%. The newer COFFE model also significantly overestimates the layout area of full adders by 13% and underestimates the layout area of multiplexers by a maximum of 60% for a prediction error variation of 73%. Such variations are particularly significant considering sensitivity analyses are not routinely performed in FPGA architectural studies. Our results suggest that such analyses are extremely important in studies that employ the minimum-width area models so the tolerance of the architectural conclusions against the prediction error variations can be quantified. Furthermore, an open-source version of the layouts of the actual FPGA building blocks should be created so their actual layout area can be used to achieve a highly accurate ranking of the implementation area of FPGA architectures built upon these layouts. Farheen Fatima Khan, Andy Gean Ye |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2017 | Measuring the Power-Constrained Performance and Energy Gap between FPGAs and Processors (Abstract Only)
Andy Gean Ye, Karthik Ganesan 0002 |
FPGA | 1 |
| 2016 | An evaluation on the accuracy of the minimum width transistor area models in ranking the layout area of FPGA architecturesabstractThis work provides an evaluation on the accuracy of the minimum width transistor area models in ranking the actual layout area of FPGA architectures. Both the original VPR area model and the new COFFE area model are compared against the actual layouts with up to 3 metal layers for the various FPGA building blocks. We found that both models have significant variations with respect to the accuracy of their predictions across the building blocks. In particular, the original VPR model overestimates the layout area of larger buffers and full adders by as much as 34% while underestimate the layout area of smaller buffers and multiplexers by as much as 59% for an overall prediction error variation of 93%. The newer COFFE model also significantly overestimates the layout area of full adders by 13% and underestimates the layout area of multiplexers by a maximum of 58% for a prediction error variation of 71%. Such variations are particularly significant considering sensitivity analyses are not routinely performed in FPGA architectural studies. Our results suggest that such analyses are extremely important in studies that employing the minimum width area models so the tolerance of the architectural conclusions against the prediction error variations can be quantified. Furthermore, an open source version of the layouts of the actual FPGA building blocks should be created, so their actual layout area can be used to achieve a highly accurate ranking of the implementation area of FPGA architectures built upon these layouts. Farheen Fatima Khan, Andy Gean Ye |
FPL | 2 |
| 2015 | Minimum jitter adaptive decision feedback equalizer for 4PAM serial linksabstractThis paper presents a minimum jitter-based adaptive decision feedback equalizer (DFE) for 4PAM serial links. For each signal level, a dedicated sign-sign least-mean-square (SS-LMS) algorithm is employed to adjust corresponding DFE tap coefficient as per data rate and the characteristics of channels. The proposed adaptive DFE is embedded in a 2 Gbps serial link over a 12-inch FR4 channel implemented in an IBM 130 nm 1.2V CMOS technology. The link is analyzed using Spectre from Cadence Design Systems with BSIM4 device models. Simulation results demonstrate that the proposed adaptive DFE is capable of opening completely closed data eyes. Alaa R. Al-Taee, Fei Yuan 0005, Andy Gean Ye |
ISCAS | 3 |
| 2014 | A new adaptive Decision Feedback Equalizer using hexagon eye-opening monitor for multi Gbps data linksabstractThis paper presents an adaptive Decision-Feedback Equalizer ADFE utilizing a proposed hexagon eye-opening monitor for multi Gbps serial links. The adaptation process of proposed ADFE depends on the error signals which are delivered from error detection unit EDU. The EDU employs a hexagon eye-opening monitor H-EOM to detect the violations of the received data signals after the comparison with three threshold voltage levels at two sampling points. The extracted error signals are then conveyed to the input of an adaptive engine. The adaptive engine updates the feedback tap coefficients of the DFE automatically based on these error signals. The examination of the comparison of the proposed DFE architect with the adaptive DFE architect employing a rectangular eye-opening monitor R-EOM shows that the proposed architect obtained better performance for providing convergence time and voltage, and identifying jitter violations. The effectiveness of the proposed architect is validated using the simulation results of a serial link designed in an IBM 130 nm 1.2V CMOS technology. Alaa R. Al-Taee, Fei Yuan 0005, Andy Gean Ye |
ISCAS | 3 |
| 2014 | New 2-D Eye-Opening Monitor for Gb/s Serial LinksabstractThis paper presents a new 2-D on-chip eye-opening monitor (EOM) for Gb/s serial links. A comprehensive review of the state-of-the-art of on-chip EOMs is provided and their pros and cons are investigated. A new hexagon 2-D EOM that outperforms the widely used rectangular 2-D EOMs is introduced and the implementation details are presented. The effectiveness of the proposed EOM is evaluated by embedding it in a serial link implemented in an IBM 130 nm 1.2 V CMOS technology. For the purpose of comparison, a rectangular 2-D EOM is also included in the same data link. The data link with a variable channel length and attenuation is analyzed using Spectre from Cadence Design Systems with BSIM four device models. Simulation results of the data link demonstrate that the proposed EOM outperforms the rectangular EOM by providing a tightened control of data jitter at the edge of data eyes and by eliminating unnecessary errors flagged by the rectangular EOM. Alaa R. Al-Taee, Fei Yuan 0005, Andy Gean Ye, Saman Sadr |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | Analysis and architecture design of scalable fractional motion estimation for H.264 encoding
Jasmina Vasiljevic, Andy Gean Ye |
Integr. | 2 |
| 2011 | VPR 5.0: FPGA CAD and architecture exploration tools with single-driver routing, heterogeneity and process scalingabstractThe VPR toolset has been widely used in FPGA architecture and CAD research, but has not evolved over the past decade. This article describes and illustrates the use of a new version of the toolset that includes four new features: first, it supports a broad range of single-driver routing architectures, which have superior architectural and electrical properties over the prior multidriver approach (and which is now employed in the majority of FPGAs sold). Second, it can now model, for placement and routing a heterogeneous selection of hard logic blocks. This is a key (but not final) step toward the incluion of blocks such as memory and multipliers. Third, we provide optimized electrical models for a wide range of architectures in different process technologies, including a range of area-delay trade-offs for each single architecture. Finally, to maintain robustness and support future development the release includes a set of regression tests for the software. To illustrate the use of the new features, we explore several architectural issues: the FPGA area efficiency versus logic block granularity, the effect of single-driver routing, and a simple use of the heterogeneity to explore the impact of hard multipliers on wiring track count. Jason Luu, Ian Kuon, Peter Jamieson, Ted Campbell, Andy Gean Ye, Wei Mark Fang, Kenneth B. Kent, Jonathan Rose |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2010 | The effect of multi-bit based connections on the area efficiency of FPGAs utilizing unidirectional routing resourcesabstractField Programmable Gate Arrays (FPGAs) are increasingly being used to implement large datapath-oriented applications that are designed to process multiple-bit wide data. Studies have shown that the regularity of these multi-bit signals can be effectively exploited to reduce the implementation area of datapath circuits on FPGAs that employ the traditional bidirectional routing. Most of modern FPGAs, however, employ unidirectional routing tracks which are more area and delay efficient. No study has investigated the design of multi-bit routing architectures to effectively transport multiple-bit wide signals using unidirectional routing tracks. This paper presents such an investigation of architectures which employ multi-bit connections and unidirectional routing resources to exploit datapath regularity. It is experimentally shown that unidirectional multi-bit routing architectures are 8.6% more area efficient than the conventional routing architecture. This paper also determines the most area efficient proportion of multi-bit routing tracks. Omesh Mutukuda, Andy Gean Ye, Gul N. Khan |
FPT | 2 |
| 2010 | Using the Minimum Set of Input Combinations to Minimize the Area of Local Routing Networks in Logic Clusters Containing Logically Equivalent I/Os in FPGAsabstractMapping digital circuits onto field-programmable gate arrays (FPGAs) usually consists of two steps. First, circuits are mapped into look-up tables (LUTs). Then, LUTs are mapped onto physical resources. The configuration of LUTs is usually determined during the first step and remains unchanged throughout the second. In this paper, we demonstrate that by reconfiguring LUTs during the second step, one can increase the flexibility of FPGA routing resources. This increase in flexibility can then be used to reduce the implementation area of FPGAs. In particular, it is shown that, for a logic cluster with$ I$inputs and$ N$$ k$-input LUTs, a set of$N\times k\quad (I+N-k+1):1$multiplexers can be used to connect logic cluster inputs to LUT inputs while maintaining logic equivalency among the logic cluster I/Os. The multiplexers (called a local routing network) are shown to be the minimum required to maintain logic equivalency. Comparing to the previous design, which employs a fully connected local routing network, the proposed design can reduce logic cluster area by 3%–25% and can reduce a significant amount of fanouts for logic cluster inputs. Andy Gean Ye |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2009 | VPR 5.0: FPGA cad and architecture exploration tools with single-driver routing, heterogeneity and process scalingabstractThe VPR toolset [6, 7] has been widely used to perform FPGA architecture and CAD research, but has not evolved over the past decade to include many architectural features now present in modern FPGAs. This paper describes a new version of the toolset that includes four significant features: first, it now supports a broad range of single-driver routing architectures [29, 4, 16]. Single-driver routing has significantly different architectural and electrical properties from the multi-driver approach previously modelled, and is now employed in the majority of FPGAs sold. Second, the new release can now model a heterogeneous selection of hard logic blocks, which could include the hard memory and multipliers that are now ubiquitous in FPGAs. Third, we provide optimized electrical models of a wide range of architectures in different process technologies, including a range of area-delay tradeoffs for each single architecture. Prior releases of VPR did not publish even one architecture file with accurate resistance and capacitance parameters. Finally, to maintain robustness and to support future development the release includes a set of regression tests to check functionality and quality of result of the output of the tools. Jason Luu, Ian Kuon, Peter Jamieson, Ted Campbell, Andy Gean Ye, Wei Mark Fang, Jonathan Rose |
FPGA | 5 |
| 2006 | Using Bus-Based Connections to Improve Field-Programmable Gate-Array Density for Implementing Datapath CircuitsabstractAs the logic capacity of field-programmable gate arrays (FPGAs) increases, they are increasingly being used to implement large arithmetic-intensive applications, which often contain a large proportion of datapath circuits. Since datapath circuits usually consist of regularly structured components (called bit-slices) which are connected together by regularly structured signals (called buses), it is possible to utilize datapath regularity in order to achieve significant area savings through FPGA architectural innovations. This paper describes such an FPGA routing architecture, called the multibit routing architecture, which employs bus-based connections in order to exploit datapath regularity. It is experimentally shown that, compared to conventional FPGA routing architectures, the multibit routing architecture can achieve 14% routing area reduction for implementing datapath circuits, which represents an overall FPGA area savings of 10%. This paper also empirically determines the best values of several important architectural parameters for the new routing architecture including the most area efficient granularity values and the most area efficient proportion of bus-based connections. Andy Gean Ye, Jonathan Rose |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2005 | Using bus-based connections to improve field-programmable gate array density for implementing datapath circuitsabstractAbstract—As the logic capacity of field-programmable gate arrays (FPGAs) increases, they are increasingly being used to implement large arithmetic-intensive applications, which often contain a large proportion of datapath circuits. Since datapath circuits usually consist of regularly structured components (called bitslices) which are connected together by regularly structured signals (called buses), it is possible to utilize datapath regularity in order to achieve significant area savings through FPGA architectural innovations. This paper describes such an FPGA routing architecture, called the multibit routing architecture, which employs busbased connections in order to exploit datapath regularity. It is experimentally shown that, compared to conventional FPGA routing architectures, the multibit routing architecture can achieve 14% routing area reduction for implementing datapath circuits, which represents an overall FPGA area savings of 10%. This paper also empirically determines the best values of several important architectural parameters for the new routing architecture including the most area efficient granularity values and the most area efficient proportion of bus-based connections. Index Terms—Area efficiency, datapath regularity, field-programmable gate arrays (FPGAs), reconfigurable fabric, routing architecture. I. Andy Gean Ye, Jonathan Rose |
FPGA | 1 |
| 2005 | Measuring and Utilizing the Correlation Between Signal Connectivity and Signal Positioning for FPGAs Containing Multi-Bit Building BlocksabstractAs the logic capacity of FPGA increases, there has been a corresponding increase in the variety of FPGA building blocks. From a mere collection of the conventional logic blocks, FPGAs now can include digital signal processors, multipliers, multi-bit addressable memory cells, and even processor cores; and one of the common characteristics of these new building blocks is their multi-bit design, where each block is designed specifically to process several bits of data at a time. This multi-bit processing paradigm is significantly different from the single-bit processing design of the conventional FPGA logic blocks; and it creates differentiation in signals through its bussed structures. Consequently, this paper examines the correlation between the positions of the signals in buses and the connectivity of these signals. Based on the correlation measurements, a multi-bit routing architecture is then proposed along with its routing tool. It is experimentally shown that, comparing to the conventional routing architectures, the multi-bit architecture requires 12% less area to implement; and in particular, it needs 27% less routing switches to connect its multi-bit blocks to their routing tracks, and 18% less configuration memory to store the configuration information. Andy Gean Ye, Jonathan Rose |
FPL | 1 |
| 2004 | Using multi-bit logic blocks and automated packing to improve field-programmable gate array density for implementing datapath circuitsabstractAs the logic capacity of field-programmable gate arrays (FPGAs) increases, they are being increasingly used to implement large arithmetic-intensive applications, which often contain a large proportion of datapath circuits. Since datapath circuits usually consist of regularly structured components, called bit-slices, it is possible to utilize datapath regularity in order to achieve significant area savings through FPGA architectural innovations. This work describes such an FPGA logic block architecture, called a multi-bit logic block, which employs configuration memory sharing to exploit datapath regularity. It is experimentally shown that, comparing to conventional FPGA logic blocks, the multi-bit logic blocks can achieve 18% to 26% logic block area reduction for implementing datapath circuits, which represents an overall FPGA area saving of 5% to 13%. A packing algorithm for the multi-bit logic block architecture is also proposed in this paper; and it is used to empirically find the best values for several important architectural parameters of the new architecture, including the most area efficient granularity values and the most area efficient amount of configuration memory sharing. Andy Gean Ye, Jonathan Rose |
FPT | 1 |
| 2002 | Synthesizing datapath circuits for FPGAs with emphasis on area minimizationabstractLarge circuits, whether they are arithmetic, digital signal processing, switching, or processors, typically contain a greater portion of highly regular datapath logic. Datapath synthesis algorithms preserve these regular structures, so they can be exploited by packing, placement, and routing tools for speed or density. Typical datapath synthesis algorithms, however, sacrifice area to gain regularity. Current algorithms can have as much as 30% to 40% area inflation when compared with traditional flat synthesis algorithms. This paper describes a datapath synthesis algorithm with very low area overhead, which is an enhancement to the module compaction algorithm. We propose two word-level optimizations - multiplexer tree collapsing and operation reordering. They reduce the area inflation to 3%-8% as compared with flat synthesis. Our synthesis results also retain significant amount of regularity from the original designs. Andy Gean Ye, Jonathan Rose, David M. Lewis |
FPT | 1 |
| 1999 | Procedural Texture Mapping on FPGAsabstractProcedural textures can be effectively used to enhance the visual realism of computer rendered images. Procedural tex-tures can provide higher realism for 3-D objects than tradi-tional hardware texture mapping methods which use mem-ory to store 2-D texture images. This paper proposes a new method of hardware texture mapping in which texture im-ages are synthesized using FPGAs. This method is very efficient for texture mapping procedural textures of more than two input variables. By synthesizing these textures on the fly, the large amount of memory required to store their multidimensional texture images is eliminated, making tex-ture mapping of 3-D textures and parameterized textures feasible in hardware. This paper shows that using FPGAs, procedural textures can be synthesized at high speed, with a small hardware cost. Data on the performance and the hardware cost of synthesizing procedural textures in FP-GAS are presented. This paper also presents, the FPGA implementations of two Perlin noise based 3-D procedural textures. Andy Gean Ye, David M. Lewis |
FPGA | 1 |