EDBT 2026 Demo / reviewers in the wild / expert
Ramprasath Srinivasa Gopalakrishnan
dblp:318/2095 · also Ramprasath S 0001, S. Ramprasath 0001
· DBLP profile ↗
13ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0001-9487-9340ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 7 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Reinforcing the Connection between Analog Design and EDA (Invited Paper)abstractBuilding upon recent advances in analog electronic design automation (EDA), this paper discusses directions for reinforcing the connection between design and EDA, in order to develop solutions that are meaningful to designers. Two aspects, both related to bridging the gap between EDA and designers, are highlighted. The first discusses the use of test structures to generate meaningful characterized data to aid design automation, specifically understanding the impact of random, correlated, and systematic variations on the design of matched structures. Results on a recent test chip that analyzes these variations and their impact on EDA design choices will be presented. The second illustrates a design testcase that applies analog EDA techniques, using the ALIGN layout engine, to design an RF MIMO receiver, and describes how this experience has helped both in advancing the state of analog EDA and in building circuits with enhanced designer productivity. Kishor Kunal, Meghna Madhusudan, Jitesh Poojary, Ramprasath Srinivasa Gopalakrishnan, Arvind K. Sharma, Ramesh Harjani, Sachin S. Sapatnekar |
ASPDAC | 4 |
| 2024 | Automated synthesis of mixed-signal ML inference hardware under accuracy constraintsabstractDue to the inherent error-tolerance of machine learning (ML) algorithms, many parts of the inference computation can be performed with adequate accuracy and low power under relatively low precision. Early approaches have used digital approximate computing methods to explore this space. Recent approaches using analog-based operations achieve power-efficient computation at moderate precision. This work proposes a mixed-signal optimization (MiSO) approach that optimally blends analog and digital computation for ML inference. Based on accuracy and power models, an integer linear programming formulation is used to optimize design metrics of analog/digital implementations. The efficacy of the method is demonstrated on multiple ML architectures. Kishor Kunal, Jitesh Poojary, Ramprasath Srinivasa Gopalakrishnan, Ramesh Harjani, Sachin S. Sapatnekar |
ASPDAC | 3 |
| 2023 | A Multicore GNN Training AcceleratorabstractGraph neural networks (GNN) are vital for analytics on real-world problems with graph models. This work develops a multicore GNN training accelerator and develops multicore-specific optimizations for superior performance. It uses enhanced multicore-specific dynamic caching to circumvent the costs of irregular DRAM access patterns of graph-structured data. A novel feature vector segmentation approach is used to maximize on-chip data reuse with high on-chip computation per memory access, reducing data access latency, using a machine learning model for optimal performance. The work presents a major advance over prior FPGA/ASIC GNN accelerators by handling significantly larger datasets (with up to 8.6M vertices) on a variety of GNN models. On average, training speedup of 17× and energy efficiency improvement of 322× is achieved over DGL on a GPU; a speedup of 14× with 268× lower energy is shown over GPU-based GNNAdvisor; and 11× and 24× speedups are obtained over ASIC-based Rubik and FPGA-based GraphACT. Sudipta Mondal, Ramprasath Srinivasa Gopalakrishnan, Ziqing Zeng, Kishor Kunal, Sachin S. Sapatnekar |
ISLPED | 2 |
| 2023 | A Unified Engine for Accelerating GNN Weighting/Aggregation Operations, With Efficient Load Balancing and Graph-Specific CachingabstractGraph neural networks (GNNs) analysis engines are vital for real-world problems that use large graph models. Challenges for a GNN hardware platform include the ability to 1) host a variety of GNNs; 2) handle high sparsity in input vertex feature vectors and the graph adjacency matrix and the accompanying random memory access patterns; and 3) maintain load-balanced computation in the face of uneven workloads, induced by high sparsity and power-law vertex degree distributions. This article proposes GNNIE, an accelerator designed to run a broad range of GNNs. It tackles workload imbalance by 1) splitting vertex feature operands into blocks; 2) reordering and redistributing computations; and 3) using a novel flexible MAC architecture. It adopts a graph-specific, degree-aware caching policy that is well suited to real-world graph characteristics. The policy enhances on-chip data reuse and avoids random memory access to DRAM. GNNIE achieves average speedups of$7197\times $over a CPU and$17.81\times $over a GPU over multiple datasets on graph attention networks (GATs), graph convolutional networks (GCNs), GraphSAGE, GINConv, and DiffPool. Compared to prior approaches, GNNIE achieves an average speedup of$5\times $over HyGCN (which cannot implement GATs) for GCN, GraphSAGE, and GINConv. GNNIE achieves an average speedup of$1.3\times $over AWB-GCN (which runs only GCNs), despite using$3.4\times $fewer processing units. Sudipta Mondal, Susmita Dey Manasi, Kishor Kunal, Ramprasath Srinivasa Gopalakrishnan, Ziqing Zeng, Sachin S. Sapatnekar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | A Generalized Methodology for Well Island Generation and Well-tap Insertion in Analog/Mixed-signal LayoutsabstractWell island generation and well tap placement is an important problem in analog/mixed-signal (AMS) circuits. Well taps can only prevent latchups within a certain radius of influence within a well island, and hence must be appropriately inserted to cover all devices. However, existing automated AMS layout paradigms typically defer the insertion of well taps and creation of well islands to a post-processing step after placement. This alters the placement, resulting in increased area and wire length, as well as circuit performance degradation. Therefore, there is a strong need for a solution that generates well islands and inserts well taps during placement so the placer can account for well overheads in optimizing placement metrics. In this work, we propose a modular solution using a graph-based optimization scheme that can be used within multiple placement paradigms with minimal intrusion. We demonstrate the integration of this scheme into stochastic, analytical, and designer-driven row-based placement. The method is demonstrated in advanced FinFET technologies. Layouts generated using this scheme show better area, wire length, and performance metrics at the cost of a marginal runtime degradation when compared to the post-processing approach. Using our scheme, there is an average improvement of 3% and 4% and a maximum improvement of 23% and 11% in area and wirelength, respectively, of layouts of various classes of AMS circuits at the cost of 17% average and 29% maximum increase in total runtime. Ramprasath Srinivasa Gopalakrishnan, Meghna Madhusudan, Arvind K. Sharma, Jitesh Poojary, Soner Yaldiz, Ramesh Harjani, Steven M. Burns, Sachin S. Sapatnekar |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2022 | GNNIE: GNN inference engine with load-balancing and graph-specific cachingabstractGraph neural networks (GNN) inferencing involves weighting vertex feature vectors, followed by aggregating weighted vectors over a vertex neighborhood. High and variable sparsity in the input vertex feature vectors, and high sparsity and power-law degree distributions in the adjacency matrix, can lead to (a) unbalanced loads and (b) inefficient random memory accesses. GNNIE ensures load-balancing by splitting features into blocks, proposing a flexible MAC architecture, and employing load (re)distribution. GNNIE's novel caching scheme bypasses the high costs of random DRAM accesses. GNNIE shows high speedups over CPUs/GPUs; it is faster and runs a broader range of GNNs than existing accelerators. Sudipta Mondal, Susmita Dey Manasi, Kishor Kunal, Ramprasath Srinivasa Gopalakrishnan, Sachin S. Sapatnekar |
DAC | 4 |
| 2022 | A Charge Flow Formulation for Guiding Analog/Mixed-Signal PlacementabstractAn analog/mixed-signal designer typically performs circuit optimization, involving intensive SPICE simulations, on a schematic netlist and then sends the optimized netlist to layout. During the layout phase, it is vital to maintain symmetry requirements to avoid performance degradation due to mismatch: these constraints are usually specified using user input or by invoking an external tool. Moreover, to achieve high performance, the layout must avoid large interconnect parasitics on critical nets. Prior works that optimize parasitics during placement work with coarse metrics such as the half-perimeter wire length, but these metrics do not appropriately emphasize performance-critical nets. The novel charge flow (CF) formulation in this work addresses both symmetry detection and parasitic optimization. By leveraging schematic-level simulations, which are available “for free” from the circuit optimization step, the approach (a) alters the objective function to emphasize the reduction of parasitics on performance-critical nets, and (b) identifies symmetric elements/element groups. The effectiveness of the CF-based approach is demonstrated on a variety of circuits within a stochastic placement engine. Tonmoy Dhar, Ramprasath Srinivasa Gopalakrishnan, Jitesh Poojary, Soner Yaldiz, Steven M. Burns, Ramesh Harjani, Sachin S. Sapatnekar |
DATE | 2 |
| 2022 | Analog/Mixed-Signal Layout Optimization using Optimal Well TapsabstractWell island generation and well tap placement pose an important challenge in automated analog/mixed-signal (AMS) layout. Well taps prevent latchup within a radius of influence in a well island, and must cover all devices. Automated AMS layout flows typically perform well island generation and tap insertion as a postprocessing step after placement. However, this step is intrusive and potentially alters the placement, resulting in increased area, wire length, and performance degradation. This work develops a graph-based optimization that integrates well island generation, well tap insertion, and placement. Its efficacy is demonstrated within a stochastic placement engine. Experimental results show that this approach generates better area, wire length and performance metrics than traditional methods, at the cost of a marginal runtime degradation. Ramprasath Srinivasa Gopalakrishnan, Meghna Madhusudan, Arvind K. Sharma, Jitesh Poojary, Soner Yaldiz, Ramesh Harjani, Steven M. Burns, Sachin S. Sapatnekar |
ISPD | 1 |
| 2016 | Efficient Algorithms for Discrete Gate Sizing and Threshold Voltage Assignment Based on an Accurate Analytical Statistical Yield GradientabstractIn this article, we derive a simple and accurate expression for the change in timing yield due to a change in the gate delay distribution. It is based on analytical bounds that we have derived for the moments of the circuit and path delay. Based on this, we propose computationally efficient algorithms for (1) discrete gate sizing and (2) simultaneous gate sizing and threshold voltage ( V T ) assignment so that the circuit meets a timing yield specification under parameter variations. The use of this analytical yield gradient within a gradient-based timing yield optimization algorithm results in a significant improvement in the runtime as compared to the numerical method, while achieving the same final yield. It also allows us to explore a larger search space in each iteration more efficiently, which is required in the case of simultaneous resizing and V T assignment. We also propose heuristics for resizing/changing the V T of multiple gates in each iteration. This makes it possible to optimize the timing yield for large circuits. Results on ITC ’99 benchmarks show that the proposed multinode resizing algorithm results in a significant improvement in the runtime with a marginal average area penalty and no cost to the final yield achieved. Ramprasath Srinivasa Gopalakrishnan, Vinita Vasudevan |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2016 | A Skew-Normal Canonical Model for Statistical Static Timing AnalysisabstractThe use of quadratic gate delay models and arrival times results in improved accuracies for a parameterized block-based statistical static timing analysis (SSTA). However, the computational complexity is significantly higher. As an alternative to this, we propose a canonical model based on skew-normal random variables (SN model). This model is derived from the quadratic canonical models and can consider the skewness in the gate delay distribution as well as the nonlinearity of the MAX operation. Based on conditional expectations, we derive the analytical expressions for the moments of the MAX operator and the tightness probability that can be used along with the SN canonical models. The computational complexity for both timing and criticality analysis is comparable with SSTA using linear models. There is a two to three orders of magnitude improvement in the run time as compared with the quadratic models. Results on ISCAS benchmarks show that the SN models have a lower variance error than the quadratic model, but the error in the third moment is comparable with that of the semiquadratic model. Ramprasath Srinivasa Gopalakrishnan, Madiwalar Vijaykumar, Vinita Vasudevan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | An efficient algorithm for statistical timing yield optimizationabstractStatistical timing yield optimization algorithms require computation of yield-gradient for gate resizing in every iteration. Numerical yield-gradients account for the effects of fan-in and fan-out gates, but are computationally expensive. In this paper, we formulate a more accurate analytical expression for the yield-gradient (termed effective yield-gradient) that includes these effects. Based on the statistical properties of the path delay variations, we derive a simplified expression for the effective yield gradient that is accurate and results in an improvement in the run-time. Using these simplified expressions, we also propose an algorithm for resizing multiple gates in an iteration. Results on ITC99 and ISCAS85 benchmarks show that the proposed multi-node resizing algorithm results in 83% improvement in the runtime with an average area penalty of 3% and no cost to the final yield achieved. Ramprasath Srinivasa Gopalakrishnan, Vinita Vasudevan |
DAC | 1 |
| 2014 | Statistical Criticality Computation Using the Circuit DelayabstractThe statistical nature of gate delays in current day technologies necessitates the use of measures, such as path criticality and node/edge criticality for timing optimization. Node criticalities are typically computed using the complementary path delay. An alternative approach to compute the criticality using the circuit delay has been recently proposed. In this paper, we discuss in detail, the use of circuit delay to compute node criticalities and show that the criticality thus found is not equal to the conventional measure found using complementary path delay. However, there is a monotonic relationship between them and the two measures can be used interchangeably. We derive new bounds for the global criticality and propose a pruning algorithm based on these bounds to improve the accuracy and speed of computation. The use of this pruning technique results in a significant speedup in criticality computations. We obtain an order of magnitude average speedup for ISCAS benchmarks. Ramprasath Srinivasa Gopalakrishnan, Vinita Vasudevan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2012 | On the computation of criticality in statistical timing analysisabstractDue to the statistical nature of gate delays in current day technologies, measures such as path criticality and node/edge criticality are required for timing optimization. Node criticalities are usually computed using the complementary path delay. In order to speed up computations, it has been recently proposed that the circuit delay be used instead. In this paper, we show that there is a monotonic relationship between the node criticalities computed using the circuit delay and the complementary delay. They are not equal, but they can be used interchangeably. We discuss the sources of error in this computation and propose methods for more accurate computations. We also introduce a measure that is very easy to compute and is an approximate indicator of criticality. Since it is easy to compute, it can also be used effectively for pruning the number of edges involved in criticality computations thus improving the speed of criticality computations. The speedup obtained can be as large as an order of magnitude for some of larger circuits in the ISCAS benchmarks. Ramprasath Srinivasa Gopalakrishnan, Vinita Vasudevan |
ICCAD | 1 |