EDBT 2026 Demo / reviewers in the wild / expert
Love Singhal
dblp:80/2878
· DBLP profile ↗
14ranked-venue papers
9as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 9 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Reconfigurable computing and FPGAs · 78% Electronic design automation · 22% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Reconfigurable computing and FPGAs
FPGA design flow |
0.3 | 1 | 2017 | LSC: A Large-Scale Consensus-Based Clustering Algorithm for High-Performance FPGAs · DAC 2017 |
Reconfigurable computing and FPGAs › FPGA physical design
logic block clustering |
0.3 | 1 | 2017 | LSC: A Large-Scale Consensus-Based Clustering Algorithm for High-Performance FPGAs · DAC 2017 |
Reconfigurable computing and FPGAs › FPGA architecture
adaptive logic module |
0.1 | 1 | 2017 | LSC: A Large-Scale Consensus-Based Clustering Algorithm for High-Performance FPGAs · DAC 2017 |
Reconfigurable computing and FPGAs
FPGA architecture |
0.1 | 1 | 2017 | LSC: A Large-Scale Consensus-Based Clustering Algorithm for High-Performance FPGAs · DAC 2017 |
Electronic design automation › physical design › timing optimization
delay budgeting |
0.1 | 1 | 2007 | Interconnect Criticality-Driven Delay Relaxation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 |
Electronic design automation › physical design › interconnect optimization
interconnect delay optimization |
0.1 | 1 | 2007 | Interconnect Criticality-Driven Delay Relaxation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 |
Electronic design automation
logic synthesis |
0.1 | 1 | 2007 | Interconnect Criticality-Driven Delay Relaxation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 |
Methods — techniques the papers use, named apart from their topics
greedy clustering · 0.3consensus-based clustering · 0.3integer linear programming · 0.1heuristic algorithm · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | High-Definition Routing Congestion Prediction for Large-Scale FPGAsabstractTo speed up the FPGA placement and routing closure, we propose a novel approach to predict the routing congestion map for large-scale FPGA designs at the placement stage. After reformulating the problem into an image translation task, our proposed approach leverages recent advancement in generative adversarial learning to address the task. Particularly, state-of-the-art generative adversarial networks for high-resolution image translation are used along with well-engineered features extracted from the placement stage. Unlike available approaches, our novel framework demonstrates a capability of handling large-scale FPGA designs. With its superior accuracy, our proposed approach can be incorporated into the placement engine to provide congestion prediction resulting in up to 7% reduction in routed wirelength for the most congested design in ISPD 2016 benchmark. Mohamed Baker Alawieh, Wuxi Li, Yibo Lin, Love Singhal, Mahesh A. Iyer, David Z. Pan |
ASP-DAC | 4 |
| 2019 | A shape-driven spreading algorithm using linear programming for global placementabstractIn this paper, we consider the problem of finding the global shape for placement of cells in a chip that results in minimum wirelength. Under certain assumptions, we theoretically prove that some shapes are better than others for purposes of minimizing wirelength, while ensuring that overlap-removal is a key constraint of the placer. We derive some conditions for the optimal shape and obtain a shape which is numerically close to the optimum. We also propose a linear-programming-based spreading algorithm with parameters to tune the resultant shape and derive a cost function that is better than total or maximum displacement objectives, that are traditionally used in many numerical global placers. Our new cost function also does not require explicit wirelength computation, and our spreading algorithm preserves to a large extent, the relative order among the cells placed after a numerical placer iteration. Our experimental results demonstrate that our shape-driven spreading algorithm improves wirelength, routing congestion and runtime compared to a bi-partitioning based spreading algorithm used in a state-of-the-art academic global placer for FPGAs. Shounak Dhar, Love Singhal, Mahesh A. Iyer, David Z. Pan |
ASP-DAC | 2 |
| 2019 | FPGA Accelerated FPGA PlacementabstractPlacement is one of the runtime bottlenecks in an FPGA design implementation flow, in which global placement accounts for a major portion of the runtime. In this paper, we demonstrate FPGA acceleration of wirelength gradient computation, which is an important part of modern analytical placement tools. To the best of our knowledge, this is the first work on acceleration of analytical placement on FPGAs. Our implementation uses OpenCL and leverages the FPGA's ability to support deep pipelines. We achieve an average speedup of 3.03x for wirelength gradient computation only and 2x for the entire global placement flow over a 28-threaded CPU implementation. Our placement quality is comparable to previously published works. Additionally, we can finish global placement for a design with 1 million cells and 1 million nets in less than 1 minute. Shounak Dhar, Love Singhal, Mahesh A. Iyer, David Z. Pan |
FPL | 2 |
| 2017 | LSC: A Large-Scale Consensus-Based Clustering Algorithm for High-Performance FPGAsabstractWith recent advances in Field Programmable Gate Array (FPGA) architecture and design, the robustness and scalability of design implementation tools is becoming increasingly important. In an FPGA implementation flow, the basic logic elements (BLEs) like flip-flops (FFs) and lookup tables (LUTs) are clustered into adaptive logic modules (ALMs) and Logic Array Blocks (LABs). Clustering is a key stage in the flow that determines whether a design can fit onto the target FPGA device, and whether the Quality of Results (QoR) goals are met. Traditionally, FPGA implementation tools have used greedy clustering techniques. This paper presents an innovative clustering algorithm based on a new concept of consensus building at a large scale (LSC). The LSC algorithm is designed to work with designs with millions of elements, and to the best of our knowledge, this is the first parallel clustering algorithm in the industry. In our industrial designs benchmark set using modern FPGA devices on two deep submicron technology nodes, the new clustering engine results in average improvements of 0.5% and 2.5% in maximum clock frequency (Fmax) for the two target devices. Additionally, wiring usage is improved on the average by 2.8% and 6.5% respectively. The fitting success rate of highly utilized designs is also improved significantly with the new clustering engine. Love Singhal, Mahesh A. Iyer, Saurabh N. Adya |
DAC | 1 |
| 2017 | An Effective Timing-Driven Detailed Placement Algorithm for FPGAsabstractIn this paper, we propose a new timing-driven detailed placement technique for FPGAs based on optimizing critical paths. Our approach extends well beyond the previously known critical path optimization approaches and explores a significantly larger solution space. It is also complementary to single-net based timing optimization approaches. The new algorithm models the detailed placement improvement problem as a shortest path optimization problem, and optimizes the placement of all elements in the entire timing critical path simultaneously, while minimizing the costs of adjusting the placement of adjacent non-critical elements. Experimental results on industrial circuits using a modern FPGA device show an average placement clock frequency improvement of 4.5%. Shounak Dhar, Mahesh A. Iyer, Saurabh N. Adya, Love Singhal, Nikolay Rubanov, David Z. Pan |
ISPD | 4 |
| 2016 | Detailed placement for modern FPGAs using 2D dynamic programmingabstractIn this paper, we propose a 2-dimensional dynamic programming (DP) based detailed placement algorithm for modern FPGAs for wirelength and timing optimization. By tuning a control parameter, our algorithm can perform fast heuristic or exact optimization. Our algorithm further enables us to solve the single row placement problem optimally which was not possible with the previous DP approaches, while also reducing it's complexity to Θ(p.N.2N) from the naive Θ(p.N!) (where p is the average degree of a net). Experiments on industrial-scale benchmarks show promising results. Shounak Dhar, Saurabh N. Adya, Love Singhal, Mahesh A. Iyer, David Z. Pan |
ICCAD | 3 |
| 2008 | Statistical power profile correlation for realistic thermal estimationabstractAt system level, the on-chip temperature depends both on power density and the thermal coupling with the neighboring regions. The problem of finding the right set of input power profile(s) for accurate temperature estimation has not been studied. Considering only average or peak power density may lead either to underestimation or overestimation of the thermal crisis, respectively. To provide more realistic temperature estimation, we propose to incorporate multiple power profiles. Using the proposed statistical methods to determine the closeness between the power profiles, we apply a clustering algorithm to identify few input power profiles. We incorporate them in a thermal-aware floorplanner and empirical results show that using the single input power profile (average or peak) leads to 37% degradation in critical wire delay and 20% degradation in wire length, compared to using the multiple input power profiles. Love Singhal, Sejong Oh, Elaheh Bozorgzadeh |
ASP-DAC | 1 |
| 2008 | Process variation aware system-level task allocation using stochastic ordering of delay distributionsabstractDesign variability due to within-die and die-to-die variations has potential to significantly reduce the maximum operating frequency and effective performance of the system in future process technology generations. When multiple cores in MPSoC have different delay distributions, the problem of assigning tasks to the cores become challenging. This paper targets system level task allocation to stochastically minimize the total execution time of an application on MPSoC under process variation. In this work, we first introduce stochastically optimal task allocation problem. We provide formal theorems of the optimality of the solution in simple scenarios. We extend our theoretical work for generic cases in normal distribution. The proposed techniques enable efficient computation of task allocation using non-stochastic analysis. We apply these techniques in allocating tasks in the embedded system benchmark suites on MPSoC. We show that deterministic solution for system-level task allocation on widely used benchmark topologies and distributions (normal distribution) is almost as good as the best probabilistic solution. Love Singhal, Elaheh Bozorgzadeh |
ICCAD | 1 |
| 2007 | Heterogeneous Floorplanner for FPGAabstractThe current generations of FPGA comprise of many specialized hardware cores, like embedded processors, multipliers, RAMs and FIFOs, along with the regular arrays of reconfigurable logic. On any FPGA device, these embedded cores are located at fixed locations only. This makes the task of floorplanning for the applications with heterogeneous components very difficult. Recently, some researchers have started looking into this problem of heterogeneous floorplanning on FPGA. However, all these work suffer from one fundamental flaw which affects the quality of solutions leading to higher device areas or excessively high runtime. In this paper, we propose a heterogeneous floorplanner for the FPGA, HPlan, which is fast and highly efficient in finding floorplans of variety of resources. We present a case study of a real implementation on Xilinx Virtex device. The proposed floorplanner could effectively implement the design with tight resource constraints whereas the traditional floorplanner could not find a feasible floorplan. Love Singhal, Elaheh Bozorgzadeh |
FCCM | 1 |
| 2007 | Novel Multi-Layer floorplanning for Heterogeneous FPGAsabstractThe current generations of FPGA comprise of many specialized hardware cores, like embedded processors, multipliers, RAMs and FIFOs, along with the regular arrays of reconfigurable logic. On any FPGA device, these embedded cores are located at fixed locations only. This makes the task of floorplanning for the applications with heterogeneous components very difficult. Recently, some researchers have started looking into this problem of heterogeneous floorplanning on FPGA. However, all these work suffer from a fundamental flaw which affects the quality of solutions leading to higher device areas or excessively high runtime. In [1], we propose a heterogeneous floorplanner for FPGA, HPlan, which is highly efficient in finding floorplans of variety of resources. In this paper, we extend the floorplanner to include an adaptive placer algorithm. We also perform our experiments on the MCNC benchmarks for the floorplan with random heterogeneous resource allocations. We observe that as the statistical variation in the heterogeneous resource allocations is increased, the traditional floorplanner gives an increasing area of all the benchmarks whereas the HPlan floorplanner does not. The proposed floorplanner thus provides an efficient way to handle floorplans with large variations in the heterogeneous resources. Love Singhal, Elaheh Bozorgzadeh |
FPL | 1 |
| 2007 | Interconnect Criticality-Driven Delay RelaxationabstractDue to decreasing transistor sizes and increasing clock frequency, interconnect delay is a dominant factor in achieving timing closure in deep-submicrometer designs. In field programmable gate arrays (FPGA), interconnect delay is contributed by programmable routing switches. This increases the wire delay significantly. In FPGA devices, the interconnect delay is usually more than 40% of the total delay. Techniques like wire pipelining and retiming can manage delay of timing critical wires. However, the latency of the design limits the total pipelining in the design. Therefore, new techniques are needed at synthesis stage to consider the effect of critical wires in the design. In this paper, we propose an intuitive Critical Edge Reduction (CER) algorithm, which minimizes the number of critical wires on a maximal delay- budgeting solution under fixed latency constraint. We prove that this problem is NP-hard. We provide an integer linear programming formulation of the problem and an iterative heuristic algorithm (CER). During the course of our algorithm, we introduce multiple graph problems. We give a proof of NP-hardness of one such problem, which we call max arc-cost balancing problem. In our experiments, we present an in-depth analysis of tradeoff between various maximal budgetings and critical edge minimization. We implemented our design flow using a set of MediaBench datapaths on Xilinx VirtexE FPGA devices. Using our algorithm, the Xilinx Place-and-Route tool achieved timing closure, which is, on average, 2.8 times faster than using maximum budgeting. The resulting average clock period using CER algorithm outperforms the one using the maximum budgeting by 6%. Other results show similar advantages of critical edge minimization over traditional budgeting techniques. Love Singhal, Elaheh Bozorgzadeh, David Eppstein |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2006 | Multi-layer Floorplanning on a Sequence of Reconfigurable DesignsabstractPartial dynamic reconfiguration is an emerging area in FPGA designs which is used for saving device area and cost. In order to reduce the reconfiguration overhead, two consecutive similar sub-designs should be placed in the same locations to get the maximum reuse of common components. This requires that all the future designs be considered while floorplanning for any given design. In this work, we introduce a new multi-layer sequence pair representation based floorplanner that allows overlap of static and non-static components of multiple designs and guarantees a feasible overlapping floorplan with minimal area packing. The multi-layer sequence pair is an efficient representation that helps in reducing the total floorplan runtime significantly. It also improves the design quality of the whole sequence as floorplans of all the designs are simultaneously computed. In our experiments, compared to a traditional sequential floorplanner, our floorplanner removes infeasibility in many designs, achieves an improvement of clock period by 12% on average and reduces the place and route time by as much as 3 times. It also reduces the average wirelength by 50% in the designs. Our proposed floorplanner could be used for finding high quality floorplans for applications that use partial reconfiguration Love Singhal, Elaheh Bozorgzadeh |
FPL | 1 |
| 2006 | Physically-aware exploitation of component reuse in a partially reconfigurable architectureabstractThe major drawback of partial dynamic reconfiguration is the reconfiguration delay overhead. To reduce the reconfiguration bitstream between two consecutive implementations, design components are reused. However, this incurs additional physical constraints to design which can lead to unroutability and congestion in design. In this paper, we propose a physically-aware component reuse strategy. We propose a floorplanning algorithm to support two-dimensional partial reconfiguration. The proposed floorplanning tool enables a wide design space exploration for component reuse. Key features are selection of the fixed modules, location of the fixed modules, mapping to the fixed modules, and interconnect planning between the fixed and reconfigurable modules. We implemented a sequence of dataflow graphs on Xilinx Virtex 4 devices using our tool for component reuse. When reuse is exploited, the experimental results report more than 50% reduction in the number of reconfiguration frames compared to the flow during which component reuse is not applied. Our proposed floorplan-aware matching technique (to map the modules to fixed components) can reduce the reconfiguration frames by 10% on average compared to dependency-based matching algorithm. In addition, we show that by different placement of the modules for two consecutive tasks, the variation in the number of reconfiguration frames can be between 25%-60% or it may even lead to unroutability of the circuits. The results imply that there is a need to tune the physical design tools for minimizing runtime reconfiguration delay overhead. Love Singhal, Elaheh Bozorgzadeh |
IPDPS | 1 |
| 2005 | Fast timing closure by interconnect criticality driven delay relaxationabstractDue to decreasing transistor sizes and increasing clock frequency, interconnect delay is a dominant factor in achieving timing closure in deep sub-micron designs. Techniques like wire pipelining and retiming can manage delay of timing critical wires. The latency of the system, however, limits the total pipelining in the design. New techniques are, thus, needed at synthesis stage to consider the effect of critical wires in the design. In this work, we propose a novel intuitive algorithm, critical edge reduction (CER) algorithm, which produces a maximal delay budgeting solution under fixed latency while minimizing the number of critical wires. We also present an in-depth analysis of trade-off between maximum budgeting and critical edge minimization. We implemented our design flow using a set of MediaBench data paths on Xilinx VirtexE FPGA devices. Using our algorithm, the Xilinx Place and Route tool achieved timing closure, on average, 2.8 times faster than using maximum budgeting. The resulting average clock period using CER algorithm outperforms the one using maximum budgeting by 6%. Love Singhal, Elaheh Bozorgzadeh |
ICCAD | 1 |