Mahesh A. Iyer

dblp:71/2489 · DBLP profile ↗
← Back
33ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0002-1045-0019ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 33 · 8 first-author · 12 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Accelerating Multi-agent Reinforcement Learning on Heterogeneous Platforms
Samuel Wiggins, Grace Zgheib, Mahesh A. Iyer, Viktor Prasanna 0001
Euro-Par (1)3
2026 FAME: A Framework for Accelerating Independent Multi-Agent Reinforcement Learning on Heterogeneous Platforms
abstract
Multi-Agent Reinforcement Learning (MARL) enables multiple autonomous agents to learn and act in a shared environment. Independent learning (IL) is a widely used MARL paradigm that underpins many real-world applications requiring efficient training at scale. However, accelerating IL at scale is non-trivial. Existing MARL frameworks rely on single-process execution and homogeneous hardware assumptions, limiting scalability and underutilizing modern heterogeneous platforms composed of CPUs, GPUs, and FPGAs. Addressing this gap requires new execution models that increase parallelism while preserving IL training semantics. In this work, we present FAME, a framework that distributes computation across heterogeneous hardware resources while providing flexible interfaces that allow MARL practitioners to prototype and test new IL approaches. FAME is composed of: (1) high-level APIs that simplify IL algorithm development, (2) a heterogeneous IL training protocol that supports concurrent agent training on multiple diverse devices, while maintaining algorithm-agnostic training semantics, (3) automatic hardware configuration generation that optimizes system throughput without needing users to manually fine-tune their system setup, and (4) dynamic load balancing among devices with different compute and memory characteristics. We demonstrate FAME’s capabilities using three representative IL algorithms on a heterogeneous node platform consisting of CPUs, GPUs, and FPGAs. Implementations generated using FAME achieve a geometric mean end-to-end training time speedup of 7.1 × over state-of-the-art implementations and up to 2.7 × speedup over additional highly parallel baselines developed in this work.
Samuel Wiggins, Nikunj Gupta, Grace Zgheib, Mahesh A. Iyer, Viktor Prasanna 0001
HPDC4
2025 Accelerating Independent Multi-Agent Reinforcement Learning on Multi-GPU Platforms
Samuel Wiggins, Nikunj Gupta, Grace Zgheib, Mahesh A. Iyer, Viktor Prasanna 0001
Euro-Par (3)4
2025 A Partitioning-Based CAD Flow for Interposer-Based Multi-Die FPGAs
abstract
Multi-die interposer-based FPGA architectures present several design challenges: (i) limited inter-die connectivity (interposer) resources and (ii) increased interposer delays compared to intra-die routing. Addressing these challenges is critical, as compared to intra-die routing, they directly impact the routability, routed wirelength (rWL) and maximum clock frequency (Fmax) of the design. In this paper, we present a partitioning-based CAD flow tailored for interposer-based multi-die FPGA architectures. Central to our approach is FPGAPart, the first open-source timing-driven netlist partitioner that can handle FPGA designs while addressing modern architectural constraints. We integrate FPGAPart with the open-source tool VTR 7.0. In particular, we use: (i) VTR 7.0's pre-packing solutions as clustering hints during partitioning and (ii) the FP-Growth algorithm [14] to detect frequently occurring patterns (instances) across multiple timing paths for clustering. Additionally, we introduce neighborhood influences-based cutting planes into the core ILP solver in FPGAPart, resulting in a ~38× ILP runtime speedup with < 1 % degradation in solution quality, compared to using no neighborhood influences. Compared to the default VTR 7.0, our flow achieves a geometric mean improvement of ~3% in rWL and ~3% in Fmax for a two-die configuration, with similar improvements across other configurations. Compared to hMETIS [17], METIS [18] and TritonPart [5], FPGAPart achieves improvements up to ~4% in rWL and ~7 % in Fmax for a two-die configuration, with similar improvements across other configurations.
Mahesh A. Iyer, Andrew B. Kahng, Jason Luu, Bodhisatta Pramanik, Kristofer Vorwerk, Grace Zgheib
FCCM1
2025 Double Duty: FPGA Architecture to Enable Concurrent LUT and Adder Chain Usage
abstract
Flexibility and customization are key strengths of Field-Programmable Gate Arrays (FPGAs) when compared to other computing devices. For instance, FPGAs can efficiently implement arbitrary-precision arithmetic operations, and can perform aggressive synthesis optimizations to eliminate ineffectual operations. Motivated by sparsity and mixed-precision in deep neural networks (DNNs), we investigate how to optimize the current logic block architecture to increase its arithmetic density. We find that modern FPGA logic block architectures prevent the independent use of adder chains, and instead only allow adder chain inputs to be fed by look-up table (LUT) outputs. This only allows one of the two primitives—either adders or LUTs—to be used independently in one logic element and prevents their concurrent use, hampering area optimizations. In this work, we propose the Double Duty logic block architecture to enable the concurrent use of the adders and LUTs within a logic element. Without adding expensive logic cluster inputs, we use 4 of the existing inputs to bypass the LUTs and connect directly to the adder chain inputs. We accurately model our changes at both the circuit and CAD levels using open-source FPGA development tools. Our experimental evaluation on a Stratix-10-like architecture demonstrates area reductions of 21.6% on adder-intensive circuits from the Kratos benchmarks, and 9.3% and 8.2% on the more general Koios and VTR benchmarks respectively. These area improvements come without an impact to critical path delay, demonstrating that higher density is feasible on modern FPGA architectures by adding more flexibility in how the adder chain is used. Averaged across all circuits from our three evaluated benchmark set, our Double Duty FPGA architecture improves area-delay product by 9.7%.
Junius Pun, Xilai Dai, Grace Zgheib, Mahesh A. Iyer, Andrew Boutros, Vaughn Betz, Mohamed S. Abdelfattah
FPL4
2025 ARC: A Runtime Engine for Accelerating Independent Multi-Agent Reinforcement Learning on Multi-Core Processors
abstract
Multi-Agent Reinforcement Learning (MARL) enables multiple agents to optimize individual or joint objectives in a shared environment, with applications spanning robotics, autonomous driving, and financial systems. Independent Learning (IL), a simple yet effective MARL approach, trains agents independently without modeling inter-agent communication or explicit coordination. This simplicity reduces computational requirements, making CPU platforms an attractive alternative to accelerators such as GPUs for smaller model architectures typical of IL. However, existing CPU-based MARL implementations rely on a Single-Learner training scheme, which sequentially trains agent networks and fails to utilize the full potential of multicore CPUs. This limits scalability and introduces inefficiencies, particularly for large-scale MARL systems. In this work, we present ARC, a lightweight runtime engine designed to accelerate IL training on multi-core CPU platforms. ARC introduces an Independent Multi-Learner training scheme that parallelizes agent model updates, maximizing hardware utilization and scalability, while preserving training semantics. By exploring and selecting optimal parallelization strategies tailored to the user's hardware, ARC ensures seamless acceleration without manual configuration. Through experiments on state-of-the-art IL algorithms, we demonstrate an increased end-to-end speedup of up to$28.2 \times$while exploring only 5% of the configuration space. We open-source ARC, supporting multiple algorithms and providing significant performance improvements, thereby facilitating the development of scalable MARL applications.
Samuel Wiggins, Nikunj Gupta, Grace Zgheib, Mahesh A. Iyer, Viktor Prasanna 0001
ICPADS4
2025 An Acceleration Framework for Deep Reinforcement Learning Using Heterogeneous Systems
abstract
Deep Reinforcement Learning (DRL) is vital in various AI applications. DRL algorithms comprise diverse compute primitives, which may not be simultaneously optimized using a homogeneous architecture. However, even with available heterogeneous architectures, optimizing DRL performance remains a challenge due to the complexity of design space in parallelizing DRL primitives and the variety of hardware employed in modern data centers. To address this, we introduce a framework for composing parallel DRL systems on heterogeneous platforms consisting of general-purpose processors (CPUs) and accelerators (GPUs, FPGAs). Our innovations include: 1. A general training protocol agnostic of the underlying hardware, enabling portable implementations across various processors and accelerators. 2. Efficient design exploration and automatic task placement enabling parallelization of tasks within each DRL primitive over one or multiple heterogeneous devices. 3. Incorporation of DRL-specific optimizations on runtime scheduling and resource allocation, facilitating parallelized training and enhancing the overall system performance. 4. High-level API for productive development using the framework. We showcase our framework through experimentation with three widely used DRL algorithms, DQN, DDPG, and SAC, on three heterogeneous platforms with diverse hardware characteristics and interconnections. The generated implementations outperform state-of-the-art libraries for CPU-GPU platforms by throughput improvements of up to 2×, and 1.7× higher performance portability across platforms.
Yuan Meng 0001, Mahesh A. Iyer, Viktor Prasanna 0001
IEEE Trans. Parallel Distributed Syst.2
2024 PEARL: Enabling Portable, Productive, and High-Performance Deep Reinforcement Learning using Heterogeneous Platforms
abstract
Deep Reinforcement Learning (DRL) is vital in various AI applications. DRL algorithms comprise diverse compute kernels, which may not be simultaneously optimized using a homogeneous architecture. However, even with available heterogeneous architectures, optimizing DRL performance remains a challenge due to the complexity of hardware and programming models employed in modern data centers. To address this, we introduce PEARL, a toolkit for composing parallel DRL systems on heterogeneous platforms consisting of general-purpose processors (CPUs) and accelerators (GPUs, FPGAs). Our innovations include: 1. A general training protocol agnostic of the underlying hardware, enabling portable implementations across various platforms. 2. Incorporation of DRL-specific optimizations on runtime scheduling and resource allocation, facilitating parallelized training and enhancing the overall system performance. 3. Automatic optimization of DRL task-to-device assignments through throughput estimation. 4. High-level API for productive development using the toolkit. We showcase our toolkit through experimentation with two widely used DRL algorithms, DQN and DDPG, on two diverse heterogeneous platforms. The generated implementations outperform state-of-the-art libraries for CPU-GPU platforms by up to 2.2× throughput improvements, and 2.4× higher performance portability across platforms.
Yuan Meng 0001, Michael Kinsner, Deshanand P. Singh, Mahesh A. Iyer, Viktor Prasanna 0001
CF4
2024 Better Together: Combining Analytical and Annealing Methods for FPGA Placement
abstract
Placement is a critical step in the FPGA design implementation flow that strongly impacts routability and timing closure. Recent state-of-the-art academic analytical placers have achieved impressive scalability but are limited to AMD Ultrascale-like architectures and mostly synthetic designs. On the other hand, VPR, the place and route tool within the widely used open-source Verilog-to-Routing (VTR) toolchain, can produce a legal placement for any arbitrary architecture; however, its simulated annealing placer scales poorly. Thus, there is a clear need to bring scalable, high-quality placement to realistic architectures and circuits. In this work, we develop a hybrid framework that combines the strength of a scalable flat analytical placer with the flexibility of simulated annealing techniques to adapt to various architectures and circuits, substantially improving the quality of results. We augment the state-of-theart analytical elfPlace FPGA placer as aug-elfPlace, generalizing its architecture modeling to handle real-world constraints and target different and more complete architectures. We leverage VPR’s legalization capability to integrate with external placers such as aug-elfPlace. VPR’s simulated annealing placer can further optimize the legalized placement, and VPR’s router and timing analysis can provide final quality results. By integrating wirelength-driven aug-elfPlace and VPR, our hybrid framework achieves up to 2% timing improvement with 15% reduction in routed wirelength compared to timing-driven VPR, on average across the large and heterogeneous Titan23 benchmark suite targeting an Intel Stratix-IV-like architecture.
Rachel Selina Rajarathnam, Kate Thurmer, Vaughn Betz, Mahesh A. Iyer, David Z. Pan
FPL4
2024 A Heterogeneous Acceleration System for Attention-Based Multi-Agent Reinforcement Learning
abstract
Multi-Agent Reinforcement Learning (MARL) is an emerging technology that has seen success in many AI applications. Multi-Actor-Attention-Critic (MAAC) is a state-of-the-art MARL algorithm that uses a Multi-Head Attention (MHA) mechanism to learn messages communicated among agents during the training process. Current implementations of MAAC using CPU and CPU-GPU platforms lack fine-grained parallelism among agents, sequentially executing each stage of the training loop, and their performance suffers from costly data movement involved in MHA communication learning. In this work, we develop the first high-throughput accelerator for MARL with attention-based communication on a CPU-FPGA heterogeneous system. We alleviate the limitations of existing implementations through a combination of data- and pipeline-parallel modules in our accelerator design and enable fine-grained system scheduling for exploiting concurrency among heterogeneous resources. Our design increases the overall system throughput by $4.6 \times$ and $4.1 \times$ compared to CPU and CPU-GPU implementations, respectively.
Samuel Wiggins, Yuan Meng 0001, Mahesh A. Iyer, Viktor Prasanna 0001
FPL3
2023 DREAMPlaceFPGA-PL: An Open-Source GPU-Accelerated Packer-Legalizer for Heterogeneous FPGAs
abstract
Placement plays a pivotal and strategic role in the FPGA implementation flow to allocate the physical locations of the heterogeneous instances in the design. Among the placement stages, the packing or clustering stage groups logic instances like look-up tables (LUTs) and flip-flops (FFs) that could be placed on the same site. The legalization stage determines all instances' physical site locations. With advances in FPGA architecture and technology nodes, designs contain millions of logic instances, and placement algorithms must scale accordingly. While other placement stages - global placement and detailed placement, have been accelerated using GPUs, the acceleration of packing and legalization stages on a GPU remains largely unexplored. This work presents DREAMPlaceFPGA-PL, an open-source packer-legalizer for heterogeneous FPGAs that employs GPU for acceleration. We revise the existing consensus-based parallel algorithms employed for packing and legalizing a flat placement to obtain further speedup on a GPU. Our experiments on the ISPD'2016 benchmarks demonstrate more than 2× acceleration.
Rachel Selina Rajarathnam, Zixuan Jiang, Mahesh A. Iyer, David Z. Pan
ISPD3
2022 DREAMPlaceFPGA: An Open-Source Analytical Placer for Large Scale Heterogeneous FPGAs using Deep-Learning Toolkit
abstract
Modern Field Programmable Gate Arrays (FPGAs) are large-scale heterogeneous programmable devices that enable high performance and energy efficiency. Placement is a crucial and computationally intensive step in the FPGA design flow that determines the physical locations of various heterogeneous instances in the design. Several works have employed GPUs and FPGAs to accelerate FPGA placement and have obtained significant runtime improvement. However, with these approaches, it is a non-trivial effort to develop optimized and algorithmic-specific kernels for GPU and FPGA to realize the best acceleration performance. In this work, we present DREAMPlaceFPGA, an open-source deep-learning toolkit-based accelerated placement framework for large-scale heterogeneous FPGAs. Notably, we develop new operators in our framework to handle heterogeneous resources and FPGA architecture-specific legality constraints. The proposed framework requires low development cost and provides an extensible framework to employ different placement optimizations. Our experimental results on the ISPD'2016 benchmarks show very promising results compared to prior approaches.
Rachel Selina Rajarathnam, Mohamed Baker Alawieh, Zixuan Jiang, Mahesh A. Iyer, David Z. Pan
ASP-DAC4
2020 High-Definition Routing Congestion Prediction for Large-Scale FPGAs
abstract
To speed up the FPGA placement and routing closure, we propose a novel approach to predict the routing congestion map for large-scale FPGA designs at the placement stage. After reformulating the problem into an image translation task, our proposed approach leverages recent advancement in generative adversarial learning to address the task. Particularly, state-of-the-art generative adversarial networks for high-resolution image translation are used along with well-engineered features extracted from the placement stage. Unlike available approaches, our novel framework demonstrates a capability of handling large-scale FPGA designs. With its superior accuracy, our proposed approach can be incorporated into the placement engine to provide congestion prediction resulting in up to 7% reduction in routed wirelength for the most congested design in ISPD 2016 benchmark.
Mohamed Baker Alawieh, Wuxi Li, Yibo Lin, Love Singhal, Mahesh A. Iyer, David Z. Pan
ASP-DAC5
2020 Symbiosis in Action: Reconfigurable Architectures and EDA
abstract
Spatial compute architectures, like Field Programmable Gate Arrays (FPGAs), constitute a key architectural pillar in modern heterogeneous compute platforms. Spatial architectures need a sophisticated Electronic Design Automation (EDA) compiler to optimally map and fit a user's workload/design onto the underlying spatial device. This EDA compiler not only helps users to custom-configure the spatial device but is also critically required for architectural exploration of new spatial architectures. The FPGA industry has had a long history of innovation in this symbiotic relationship between EDA and reconfigurable spatial architectures.
Mahesh A. Iyer
FPGA1
2020 Agilex™ Generation of Intel® FPGAs
abstract
This article consists only of a collection of slides from the author's conference presentation.
Ilya Ganusov, Mahesh A. Iyer, Alon Meisler
Hot Chips Symposium2
2019 A shape-driven spreading algorithm using linear programming for global placement
abstract
In this paper, we consider the problem of finding the global shape for placement of cells in a chip that results in minimum wirelength. Under certain assumptions, we theoretically prove that some shapes are better than others for purposes of minimizing wirelength, while ensuring that overlap-removal is a key constraint of the placer. We derive some conditions for the optimal shape and obtain a shape which is numerically close to the optimum. We also propose a linear-programming-based spreading algorithm with parameters to tune the resultant shape and derive a cost function that is better than total or maximum displacement objectives, that are traditionally used in many numerical global placers. Our new cost function also does not require explicit wirelength computation, and our spreading algorithm preserves to a large extent, the relative order among the cells placed after a numerical placer iteration. Our experimental results demonstrate that our shape-driven spreading algorithm improves wirelength, routing congestion and runtime compared to a bi-partitioning based spreading algorithm used in a state-of-the-art academic global placer for FPGAs.
Shounak Dhar, Love Singhal, Mahesh A. Iyer, David Z. Pan
ASP-DAC3
2019 FPGA Accelerated FPGA Placement
abstract
Placement is one of the runtime bottlenecks in an FPGA design implementation flow, in which global placement accounts for a major portion of the runtime. In this paper, we demonstrate FPGA acceleration of wirelength gradient computation, which is an important part of modern analytical placement tools. To the best of our knowledge, this is the first work on acceleration of analytical placement on FPGAs. Our implementation uses OpenCL and leverages the FPGA's ability to support deep pipelines. We achieve an average speedup of 3.03x for wirelength gradient computation only and 2x for the entire global placement flow over a 28-threaded CPU implementation. Our placement quality is comparable to previously published works. Additionally, we can finish global placement for a design with 1 million cells and 1 million nets in less than 1 minute.
Shounak Dhar, Love Singhal, Mahesh A. Iyer, David Z. Pan
FPL3
2017 LSC: A Large-Scale Consensus-Based Clustering Algorithm for High-Performance FPGAs
abstract
With recent advances in Field Programmable Gate Array (FPGA) architecture and design, the robustness and scalability of design implementation tools is becoming increasingly important. In an FPGA implementation flow, the basic logic elements (BLEs) like flip-flops (FFs) and lookup tables (LUTs) are clustered into adaptive logic modules (ALMs) and Logic Array Blocks (LABs). Clustering is a key stage in the flow that determines whether a design can fit onto the target FPGA device, and whether the Quality of Results (QoR) goals are met. Traditionally, FPGA implementation tools have used greedy clustering techniques. This paper presents an innovative clustering algorithm based on a new concept of consensus building at a large scale (LSC). The LSC algorithm is designed to work with designs with millions of elements, and to the best of our knowledge, this is the first parallel clustering algorithm in the industry. In our industrial designs benchmark set using modern FPGA devices on two deep submicron technology nodes, the new clustering engine results in average improvements of 0.5% and 2.5% in maximum clock frequency (Fmax) for the two target devices. Additionally, wiring usage is improved on the average by 2.8% and 6.5% respectively. The fitting success rate of highly utilized designs is also improved significantly with the new clustering engine.
Love Singhal, Mahesh A. Iyer, Saurabh N. Adya
DAC2
2017 An Effective Timing-Driven Detailed Placement Algorithm for FPGAs
abstract
In this paper, we propose a new timing-driven detailed placement technique for FPGAs based on optimizing critical paths. Our approach extends well beyond the previously known critical path optimization approaches and explores a significantly larger solution space. It is also complementary to single-net based timing optimization approaches. The new algorithm models the detailed placement improvement problem as a shortest path optimization problem, and optimizes the placement of all elements in the entire timing critical path simultaneously, while minimizing the costs of adjusting the placement of adjacent non-critical elements. Experimental results on industrial circuits using a modern FPGA device show an average placement clock frequency improvement of 4.5%.
Shounak Dhar, Mahesh A. Iyer, Saurabh N. Adya, Love Singhal, Nikolay Rubanov, David Z. Pan
ISPD2
2017 CAD Opportunities with Hyper-Pipelining
abstract
Hyper-pipelining is a design technique that results in significant performance and throughput improvements in latency-insensitive designs. Modern FPGA architectures like Intel's Stratix®10 feature a revolutionary register-rich HyperFlex? core fabric architecture that make it amenable for hyper-pipelining. Design implementation CAD tools can provide insights into performance bottlenecks and how hyper-pipelining can result in improved performance, that can then be implemented using well-known techniques like retiming. Retiming was first introduced as a powerful sequential design optimization technique three decades ago, yet gained limited popularity in the ASIC industry. In recent years, retiming has gained tremendous popularity in the FPGA industry. This talk will discuss why this is the case, and provide insights into some of the interesting opportunities it presents for design implementation, analysis, and verification CAD tools. Impacts of hyper-pipelining on the physical design CAD flow and timing closure will also be discussed.
Mahesh A. Iyer
ISPD1
2016 Detailed placement for modern FPGAs using 2D dynamic programming
abstract
In this paper, we propose a 2-dimensional dynamic programming (DP) based detailed placement algorithm for modern FPGAs for wirelength and timing optimization. By tuning a control parameter, our algorithm can perform fast heuristic or exact optimization. Our algorithm further enables us to solve the single row placement problem optimally which was not possible with the previous DP approaches, while also reducing it's complexity to Θ(p.N.2N) from the naive Θ(p.N!) (where p is the average degree of a net). Experiments on industrial-scale benchmarks show promising results.
Shounak Dhar, Saurabh N. Adya, Love Singhal, Mahesh A. Iyer, David Z. Pan
ICCAD4
2009 On improving optimization effectiveness in interconnect-driven physical synthesis
abstract
In modern designs, the delay of a net can vary significantly depending on its routing. This large estimation error during the pre-routing stage can often mislead the optimization of the netlist. We extend state-of-the-art interconnect-driven physical synthesis by introducing a new paradigm (namely, persistence) that relies on guaranteed net routes for the most sensitive nets while performing circuit optimization in the pre-route stage. We implemented our proposed approach in a cutting-edge industrial physical synthesis flow; this involved the automatic identification and routing of critical nets that were likely to be mispredicted, the automatic update of their routes during the subsequent pre-routing stage optimizations, and the guaranteed retention of their routes across the routing stage. Our approach achieves significant performance improvements on a suite of real-world 65nm designs, while ensuring that the impact on their routability remains negligible. Furthermore, our experimental results scale very well with design size.
Prashant Saxena, Vishal Khandelwal, Changge Qiao, Pei-Hsin Ho, J.-C. Lin, Mahesh A. Iyer
ISPD6
2003 Race: A Word-Level ATPG-Based Constraints Solver System For Smart Random Simulation
Mahesh A. Iyer
ITC1
2000 FILL and FUNI: algorithms to identify illegal states and sequentially untestable faults
abstract
In this paper, we first present an algorithm (FILL) to efficiently identify a large subset of illegal states in synchronous sequential circuits, without assuming a global reset mechanism. A second algorithm, FUNI, finds sequentially untestable faults whose detection requires some of the illegal states computed by FILL. Although based on binary decision diagrams (BDDs), FILL is able to process large circuits by using a new functional partitioning procedure. The incremental building of the set of illegal states guarantees that FILL will always obtain at least a partial solution. FUNI is a direct method that identifies untestable faults without using the exhaustive search involved in automatic test generation (ATG). Experimental results show that FUNI finds a large number of untestable faults up to several orders of magnitude faster than an ATG algorithm that targeted the faults identified by FUNI. Also, many untestable faults identified by FUNI were aborted by the test generator.
David E. Long, Mahesh A. Iyer, Miron Abramovici
ACM Trans. Design Autom. Electr. Syst.2
1999 Wavefront Technology Mapping
abstract
The wavefront technology mapping algorithm leads to a very simple and efficient implementation that elegantly decouples pattern matching and covering but circumvents that patterns have to be stored for the entire network simultaneously. This coupled with dynamic decomposition enables trade-off of many more alternatives than in conventional mapping algorithms. The wavefront algorithm maps optimally for minimal delay on directed acyclic graphs (DAGs) when a gain based delay model is used. It is optimal with respect to the arrival times on each path in the network. A special timing mode for multi-source nets allows minimization of other (non-delay) metrics as a secondary objective while maintaining delay optimality.
Leon Stok, Andrew J. Sullivan, Mahesh A. Iyer
DATE3
1999 A Robust Solution to the Timing Convergence Problem in High-Performance Design
abstract
Traditional ASIC design flows have treated logic synthesis and physical design as separate steps in the flow. A recent trend in design automation has been to integrate placement and logic synthesis operations for designs that strive for high performance. The motivation for this is ascribed to achieving timing convergence. These efforts attempt a brute-force combination of techniques from the two fields. We present an architecture for combining synthesis transforms with rough placement. There are three main contributions of this paper. First we present a system architecture that permits a clean separation of placement and synthesis issues and combines the two solutions in an elegant manner. Second, we propose a minor modification to the current ASIC design flow to enable timing convergence. Third, we use design rules for correct circuit operation to drive the placement and the synthesis components of the system. We present results for a set of high performance ASIC designs which demonstrate the practicality of our method.
Narendra V. Shenoy, Mahesh A. Iyer, Robert F. Damiano, Kevin Harer, Hi-Keung Tony Ma, Paul Thilking
ICCD2
1999 High Time For High Level ATPG
Mahesh A. Iyer
ITC1
1996 Identifying Sequential Redundancies Without Search
abstract
personal or class-room use is granted without fee provided that copies are not made or distributed for profit or commercial advantage, the copyright notice, the title of the publication and its date appear, and notice is given that copying is
Mahesh A. Iyer, David E. Long, Miron Abramovici
DAC1
1996 FIRE: a fault-independent combinational redundancy identification algorithm
abstract
FIRE is a novel Fault-Independent algorithm for combinational REdundancy identification. The algorithm is based on a simple concept that a fault which requires a conflict as a necessary condition for its detection is undetectable and hence redundant. FIRE does not use the backtracking-based exhaustive search performed by fault-oriented automatic test generation algorithms, and identifies redundant faults without any search. Our results on benchmark and real circuits indicate that we find a large number of redundancies (about 80% of the combinational redundancies in benchmark circuits), much faster than a test-generation-based approach for redundancy identification. However, FIRE is not guaranteed to identify all redundancies in a circuit.
Mahesh A. Iyer, Miron Abramovici
IEEE Trans. Very Large Scale Integr. Syst.1
1995 Identifying sequentially untestable faults using illegal states
abstract
In this paper, we first present an algorithm (FILL) which efficiently identifies a large subset of the illegal states in a synchronous sequential circuit, without assuming a global reset mechanism. A second algorithm, FUNI, finds sequentially untestable faults whose detection requires some of the illegal states computed by FlLL. Although based on binary decision diagrams (BDDs), FILL is able to process large circuits by using a new functional partitioning procedure. The incremental building of the set of illegal states guarantees that FILL mill always obtain at least a partial solution. FUNI is a direct method that identifies untestable faults without using the exhaustive search involved in automatic test generation (ATG). Experimental results show that FUNI finds a large number of untestable faults up to several orders of magnitude faster than an ATG algorithm that targeted the faults identified by FUNI, Also, many untestable faults identified by FUNI were aborted by the test generator.
David E. Long, Mahesh A. Iyer, Miron Abramovici
VTS2
1995 Energy models for delay testing
abstract
We present a new formulation of the delay testing problem as an energy minimization problem. Two important applications have motivated this work. First, it can be used to efficiently generate robust and nonrobust tests for path delay faults in scan and hold type of sequential circuits. Second, It allows the design of a special class of delay fault testable circuits, called (k,K)-circuits, that have polynomial-time test generation complexity. For the new formulation, the relationship between input and output signal states of a logic gate for an arbitrary pair of input vectors is expressed through an energy function. The minimum-energy states of this function correspond to signal values that are consistent with the gate's logic function. The function also implicitly includes the information about the potential hazards due to arbitrary delay distributions in the circuit. The energy function for the circuit is the summation of the individual gate energy functions. To derive tests for a given delay fault, this function is suitably modified such that any minimum-energy state is guaranteed to be a test. The specific modifications to the energy function depend on the type (robust or nonrobust, with or without hazards) of delay test desired. For (k, K)-circuits, we show that the energy function can be minimized in polynomial-time. For general circuits, where the problem still has an exponential complexity, the recently proposed transitive closure based test generation technique is very effective in generating tests. This approach efficiently determines a delay test or establishes that no test is possible for the given delay fault. We report experimental results on various sequential benchmark circuits (full-scan versions) showing the feasibility and practicality of the new methods.>
Srimat T. Chakradhar, Mahesh A. Iyer, Vishwani D. Agrawal
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1994 Sequentially Untestable Faults Identified Without Search ("Simple Implications Beat Exhaustive Search!")
abstract
This paper presents a novel fault-independent algorithm for identifying untestable faults in sequential circuits. The algorithm is based on a simple concept that a fault which requires an illegal combination of values as a necessary condition for its detection is untestable. It uses implications to find a subset of such faults whose detection requires conflicts on certain lines in the circuit. No global reset state is assumed and no state transition information is needed. Our fault-independent algorithm identifies untestable faults without any search as opposed to exhaustive search done by fault-oriented test generation algorithms. Results on benchmark and real circuits indicate that we find a large number of untestable faults, much faster (up to 3 orders of magnitude) than a test-generation-based algorithm that targeted the faults identified by our algorithm. Moreover, many faults identified as untestable by our approach were aborted when targeted by a sequential test generator.
Mahesh A. Iyer, Miron Abramovici
ITC1
1992 One-Pass Redundancy Identification and Removal
Miron Abramovici, Mahesh A. Iyer
ITC2