VLDB 2026 Research / reviewers in the wild / expert
Andrew B. Kahng
dblp:k/AndrewBKahng
· DBLP profile ↗
390ranked-venue papers
160as first author
51since 2021 · last 2026
0000-0002-4490-5018ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 378 · 155 first-author · 50 since 2021Software engineering, systems software and programming languages · 22 · 11 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 first-authorTheory of computation · 5Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Au-MEDAL: Adaptable Grid Router with Metal Edge Detection And Layer Integration
Andrew B. Kahng, Seokhyeong Kang, Jakang Lee, Dooseok Yoon |
ASP-DAC | 1 |
| 2026 | LSMC Meets GPU Acceleration: Scalable and High-Quality Multi-Row Detailed Placement
Andrew B. Kahng, Jason Liang, Zhiang Wang |
ISCAS | 1 |
| 2026 | Invited: Agentic AI for Physical Design R&D: Status and ProspectsabstractRecent advances in large language models (LLMs) and tool-using autonomous agents present new opportunities for accelerating research and development in physical design. Unlike earlier uses of machine learning that focused narrowly on prediction or optimization subroutines, agentic AI systems can comprehend user specifications, modify code, run EDA tools, analyze results, perform multi-step reasoning, and iteratively refine design heuristics. This paper surveys the emerging landscape of agentic AI for physical design R&D, with emphasis on (i) tool-integrated agents for algorithm evolution, debugging, and workflow automation, (ii) autonomous exploration of heuristic spaces in placement, routing, and partitioning, and (iii) interfaces between agents and traditional EDA frameworks. We analyze recent experience with multi-agent workflows and benchmark evaluation, highlighting current capabilities, limitations, and research frontiers. We conclude by articulating the long-term prospects of agentic AI as a catalyst for accelerated innovation in physical design, including autonomous algorithm discovery, continuous tool improvement, and closed-loop learning from large design corpora. Amur Ghose, Andrew B. Kahng, Sayak Kundu, Bodhisatta Pramanik |
ISPD | 2 |
| 2026 | Invited: Toward Sustainable and Transparent Benchmarking for Academic Physical Design ResearchabstractThis paper presents RosettaStone 2.0, an open benchmark translation and evaluation framework built on OpenROAD-Research [1]. RosettaStone 2.0 provides complete RTL-to-GDS reference flows for both conventional 2D designs and Pin-3D-style face-to-face (F2F) hybrid-bonded 3D designs, enabling rigorous apples-to-apples comparison across planar and three-dimensional implementation settings. The framework is integrated within OpenROAD-flowscripts (ORFS)-Research [2]; it incorporates continuous integration (CI)-based regression testing and provides a standardized evaluation pipeline based on the METRICS2.1 convention, with structured logs and reports generated by ORFS-Research. To support transparent and reproducible research, RosettaStone 2.0 further provides a community-facing leaderboard, which is governed by verified pull requests and enforced through Developer Certificate of Origin (DCO) compliance. Andrew B. Kahng, Zhiang Wang, Zhiyu Zheng |
ISPD | 2 |
| 2026 | Invited: The Many PD Faces of Professor Jens LienigabstractThe 2026 ISPD Lifetime Achievement Award honors Professor Jens Lienig for his tremendous decades-long impact on physical design research, education, and community. Professor Lienig's career uniquely weaves algorithmic methodology, industry-driven foci, and pedagogical codification for future generations. It has both mirrored and driven the broadening of physical design's objective functions and notions of correctness, from geometric checks to reliability- and physics-aware methodologies. This invited paper provides a few personal perspectives to complement the excellent review by Knechtel et al. [46]. Three ''PD faces'' of Jens Lienig are highlighted: (i) early evolutionary and parallel metaheuristics for combinatorial problems in physical design; (ii) textbooks whose arc spans system context through layout practice and physical design automation, alongside landmark monographs on electromigration-aware design and 3D integration; and (iii) influence on the trajectory, research and community of ISPD. Andrew B. Kahng |
ISPD | 1 |
| 2026 | Invited: Post-Placement Buffering and Sizing ContestabstractThe ISPD 2026 Contest [22] challenges participants to develop post-detailed placement buffering and sizing tools that optimize timing and fix electrical rule check (ERC) violations under real-world constraints. Unlike prior contests, this contest emphasizes practical physical design challenges including fixed macros and I/Os, power delivery network (PDN) blockages, soft placement blockages, and fixed routing resources. The contest provides eight public benchmarks and four hidden benchmarks, with a range from 15K to 1.4M instances, in the ASAP7 7nm technology node [4] with multi-threshold voltage cell libraries. Evaluation is performed using the open-source OpenROAD infrastructure, with scoring based on timing (total negative slack), power (dynamic and leakage) and penalties for ERC violations, displacement, routing congestion and runtime. This paper describes the contest problem formulation, benchmarks, evaluation methodology, a review of related contests and a two-year roadmap for continuation in the ISPD 2027 Contest. Andrew B. Kahng, Seokhyeong Kang, Sayak Kundu, Yiting Liu 0002, Davit Markarian, Seonghyeon Park, Zhiang Wang |
ISPD | 1 |
| 2026 | An Updated Assessment of Reinforcement Learning for Macro PlacementabstractWe provide an improved assessment of Google Brain’s deep reinforcement learning approach to macro placement [29] and its updated Circuit Training (CT) implementation in GitHub [53]. A stronger simulated annealing (SA) baseline leverages the “go-with-the-winners” metaheuristic [3] and a multi-threading implementation. We develop and release new public benchmarks in sub-10nm technology: LEF/DEF for Google’s 7nm TSMC Ariane protobuf and scaled variants, as well as testcases implemented in the open-source ASAP7 7nm research enablement. We evaluate from-scratch training and fine-tuning results for the latest “AlphaChip” release of Circuit Training, alongside multiple alternative macro placers. We also study the recently-published pre-training guidance in [53]. A commercial place-and-route tool is used to provide “true reward” post-route power, performance and area metrics. All data, evaluation flows and related scripts are publicly available in theMacroPlacementGitHub repository [63]. Our study affords insights into reproducibility and reporting in the research literature, and points out still-missing confirmations (e.g., of CT’s scalability and pre-training methodology) that remain open questions for the research community. Chung-Kuan Cheng, Andrew B. Kahng, Sayak Kundu, Yucheng Wang 0016, Zhiang Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | ChipletPart: Cost-Aware Partitioning for 2.5D SystemsabstractIndustry adoption of chiplets has been growing as chiplets are a cost-effective option for making large, high-performance systems. Consequently, partitioning large systems into chiplets is increasingly important. In this work, we introduce ChipletPart —a cost-driven 2.5D system partitioner that addresses the unique constraints of chiplet systems, including complex objective functions, limited reach of inter-chiplet I/O transceivers, and the assignment of heterogeneous manufacturing technologies to different chiplets. ChipletPart integrates a sophisticated chiplet cost model with a genetic algorithm (GA)-based technology assignment and partitioning methodology, along with a simulated annealing (SA)-based chiplet floorplanner. Our results show that ChipletPart : (i) reduces chiplet cost by up to 58% (20% geometric mean) compared to state-of-the-art min-cut partitioners, which often yield floorplan-infeasible solutions; (ii) generates partitions with up to 47% (6% geometric mean) lower cost compared to the prior work Floorplet ; (iii) reduces chiplet cost up to 48% (30% geometric mean) compared to Chipletizer , while consistently producing I/O-feasible chiplet solutions across all testcases; and (iv) for the testcases we study, heterogeneous integration reduces cost by up to 43% (15% geometric mean) compared to homogeneous implementations. Additionally, we explore Bayesian optimization (BO) for finding low cost and floorplan-feasible chiplet solutions with technology assignments. On some testcases, our BO framework achieves better system cost (up to 5.3% improvement) with higher runtime overhead (up to 4×) compared to our GA-based framework. We also present case studies that show how changes in packaging and inter-chiplet signaling technologies can affect partitioning solutions. Finally, ChipletPart , the underlying chiplet cost model, and our chiplet testcase generator are available as open-source tools for the community. Alexander Graening, Puneet Gupta 0001, Andrew B. Kahng, Bodhisatta Pramanik, Zhiang Wang |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2025 | Use Cases and Deployment of ML in IC Physical DesignabstractML for IC physical design must be deployed in order to have business impacts. However, deployment in production must navigate many practical considerations, including choice of targets, skillsets and infrastructure, expectations and resources, data, and "MLOps". Furthermore, usage of ML is not the same as IC design practice and capability. In this invited paper, we give perspectives on basic strategies for selecting applications and pursuing deployment for ML in IC physical design. Example aspects include checklists for data and ML models, evaluation of model performance and progress on the path to deployment, the shifting landscape of MLOps, and challenges of "LLM-ability". Amur Ghose, Andrew B. Kahng, Sayak Kundu, Yiting Liu 0002, Bodhisatta Pramanik, Zhiang Wang, Dooseok Yoon |
ASP-DAC | 2 |
| 2025 | A Partitioning-Based CAD Flow for Interposer-Based Multi-Die FPGAsabstractMulti-die interposer-based FPGA architectures present several design challenges: (i) limited inter-die connectivity (interposer) resources and (ii) increased interposer delays compared to intra-die routing. Addressing these challenges is critical, as compared to intra-die routing, they directly impact the routability, routed wirelength (rWL) and maximum clock frequency (Fmax) of the design. In this paper, we present a partitioning-based CAD flow tailored for interposer-based multi-die FPGA architectures. Central to our approach is FPGAPart, the first open-source timing-driven netlist partitioner that can handle FPGA designs while addressing modern architectural constraints. We integrate FPGAPart with the open-source tool VTR 7.0. In particular, we use: (i) VTR 7.0's pre-packing solutions as clustering hints during partitioning and (ii) the FP-Growth algorithm [14] to detect frequently occurring patterns (instances) across multiple timing paths for clustering. Additionally, we introduce neighborhood influences-based cutting planes into the core ILP solver in FPGAPart, resulting in a ~38× ILP runtime speedup with < 1 % degradation in solution quality, compared to using no neighborhood influences. Compared to the default VTR 7.0, our flow achieves a geometric mean improvement of ~3% in rWL and ~3% in Fmax for a two-die configuration, with similar improvements across other configurations. Compared to hMETIS [17], METIS [18] and TritonPart [5], FPGAPart achieves improvements up to ~4% in rWL and ~7 % in Fmax for a two-die configuration, with similar improvements across other configurations. Mahesh A. Iyer, Andrew B. Kahng, Jason Luu, Bodhisatta Pramanik, Kristofer Vorwerk, Grace Zgheib |
FCCM | 2 |
| 2025 | SO3-Cell: Standard Cell Layout Automation Framework for Simultaneous Optimization of Topology, Placement, and RoutingabstractWe propose SO3-Cell, the first automatic standard cell layout generation framework that optimizes three key steps simultaneously using Mixed-Integer Linear Programming (MILP). SO3-Cell simultaneously performs circuit topology optimization, transistor placement, and internal cell routing to achieve an optimized layout solution. Our optimization objective is to minimize metal usage while enhancing cell layout flexibility within a given area.We introduce design space pruning techniques to mitigate the complexity of larger designs, such as a full adder, a reset flip-flop (FF), and a 2-bit FF. We successfully generate a layout for a 44-transistor 2-bit FF within 25,862 seconds, demonstrating the scalability and robustness of the SO3-Cell framework. We evaluate the block-level PPA impact of the proposed cell-layout improvements, demonstrating a 35.0% reduction in power, a 2.2% increase in frequency, and a 31.1% reduction in area. Chung-Kuan Cheng, Andrew B. Kahng, Byeonggon Kang, Seokhyeong Kang, Jakang Lee, Bill Lin 0001 |
ICCAD | 2 |
| 2025 | Invited: IEEE DATC RDF-2025: Enabling an EDA Research EcosystemabstractOver the past year, IEEE CEDA DATC has continued to improve the DATC Robust Design Flow (RDF) while continuing to expand initiatives that advance open infrastructures and culture changes, serving the global community of EDA researchers and users. This invited paper focuses on three highlights: (1) establishment of an accessible, "contrib-like" GitHub resource that provides a more accessible environment for OpenROAD- and OpenROAD-flow-scripts-based research works; (2) the first-ever permission mechanism and benchmarking results for a commercial EDA P&R tool, published with permissions developed with the tool vendor (Siemens EDA); and (3) efforts that support a nascent "ML EDA Commons". The paper also provides brief reviews of the past year’s RDF developments and roadmap updates. Vidya A. Chhabria, Amur Ghose, Vikram Gopalakrishnan, Andrew B. Kahng, Sayak Kundu, Yiting Liu 0002, Zhiang Wang, Bing-Yue Wu |
ICCAD | 4 |
| 2025 | Invited: Toward an ML EDA Commons: Establishing Standards, Accessibility, and Reproducibility in ML-driven EDA ResearchabstractMachine learning (ML) is transforming electronic design automation (EDA), offering innovative solutions for designing and optimizing integrated circuits (ICs). However, the field faces significant challenges in standardization, accessibility, and reproducibility, limiting the impact of ML-driven EDA (ML EDA) research. To address these barriers, this paper presents a vision for an ML EDA Commons, a collaborative open ecosystem designed to unify the community and drive progress through establishing standards, shared resources, and stakeholder-based governance. The ML EDA Commons focuses on three objectives: (1) Maturing existing EDA infrastructure to support ML EDA research; (2) Establishing standards for benchmarks, metrics, and data quality and formats for consistent evaluation via governance that includes key stakeholders; and (3) Improving accessibility and reproducibility by providing open datasets, tools, models, and workflows with cloud computing resources, to lower barriers to ML EDA research and promote robust research practices via artifact evaluations, canonical evaluators, and integration pipelines. Inspired by successes of ML and MLCommons, the ML EDA Commons aims to catalyze transparency and sustainability in ML EDA research. Vidya A. Chhabria, Jiang Hu 0001, Andrew B. Kahng, Sachin S. Sapatnekar |
ISPD | 3 |
| 2025 | DG-RePlAce: A Dataflow-Driven GPU-Accelerated Analytical Global Placement Framework for Machine Learning AcceleratorsabstractGlobal placement (GP) is a fundamental step in VLSI physical design. The wide use of 2-D processing element (PE) arrays in machine learning accelerators poses new challenges of scalability and quality of results (QoR) for state-of-the-art academic global placers. In this work, we develop DG-RePlAce, a new and fast GPU-accelerated GP framework built on top of the OpenROAD infrastructure, which exploits the inherent dataflow and datapath structures of machine learning accelerators. Experimental results with a variety of machine learning accelerators using a commercial 12-nm enablement show that, compared with RePlAce (DREAMPlace), our approach achieves an average reduction in routed wirelength by$10\%~(7\%)$and total negative slack (TNS) by$31\%~(34\%)$, with faster GP and on-par total runtimes relative to DREAMPlace. Empirical studies on the TILOS MacroPlacement Benchmarks further demonstrate that post-route improvements over RePlAce and DREAMPlace may reach beyond the motivating application to machine learning accelerators. Andrew B. Kahng, Zhiang Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | Performance Analysis of CNN Inference/Training with Convolution and Non-Convolution Operations on ASIC AcceleratorsabstractToday’s performance analysis frameworks for deep learning accelerators suffer from two significant limitations. First, although modern convolutional neural networks (CNNs) consist of many types of layers other than convolution, especially during training, these frameworks largely focus on convolution layers only. Second, these frameworks are generally targeted towards inference and lack support for training operations. This work proposes a novel open-source performance analysis framework, SimDIT, for general ASIC-based systolic hardware accelerator platforms. The modeling effort of SimDIT comprehensively covers convolution and non-convolution operations of both CNN inference and training on a highly parameterizable hardware substrate. SimDIT is integrated with a backend silicon implementation flow and provides detailed end-to-end performance statistics (i.e., data access cost, cycle counts, energy, and power) for executing CNN inference and training workloads. SimDIT-enabled performance analysis reveals that on a 64×64 processing array, non-convolution operations constitute 59.5% of total runtime for ResNet-50 training workload. In addition, by optimally distributing available off-chip DRAM bandwidth and on-chip SRAM resources, SimDIT achieves 18× performance improvement over a generic static resource allocation for ResNet-50 inference. Hadi Esmaeilzadeh, Soroush Ghodrati, Andrew B. Kahng, Sean Kinzer, Susmita Dey Manasi, Sachin S. Sapatnekar, Zhiang Wang |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2024 | NN-Steiner: A Mixed Neural-Algorithmic Approach for the Rectilinear Steiner Minimum Tree ProblemabstractRecent years have witnessed rapid advances in the use of neural networks to solve combinatorial optimization problems. Nevertheless, designing the "right" neural model that can effectively handle a given optimization problem can be challenging, and often there is no theoretical understanding or justification of the resulting neural model. In this paper, we focus on the rectilinear Steiner minimum tree (RSMT) problem, which is of critical importance in IC layout design and as a result has attracted numerous heuristic approaches in the VLSI literature. Our contributions are two-fold. On the methodology front, we propose NN-Steiner which is a novel mixed neural-algorithmic framework for computing RSMTs that leverages the celebrated PTAS algorithmic framework of Arora to solve this problem (and other geometric optimization problems). Our NN-Steiner replaces key algorithmic components within Arora's PTAS by suitable neural components. In particular, NN-Steiner only needs four neural network (NN) components that are called repeatedly within an algorithmic framework. Crucially, each of the four NN components is only of bounded size independent of input size, and thus easy to train. Furthermore, as the NN component is learning a generic algorithmic step, once learned, the resulting mixed neural-algorithmic framework generalizes to much larger instances not seen in training. Our NN-Steiner, to our best knowledge, is the first neural architecture of bounded size that has capacity to approximately solve RSMT (and variants). On the empirical front, we show how NN-Steiner can be implemented and demonstrate the effectiveness of our resulting approach, especially in terms of generalization, by comparing with state-of-the-art methods (both neural and non-neural based). Andrew B. Kahng, Robert R. Nerem, Yusu Wang 0001, Chien-Yi Yang |
AAAI | 1 |
| 2024 | PPA-Relevant Clustering-Driven Placement for Large-Scale VLSI DesignsabstractToday's place-and-route (P&R) flows are increasingly challenged by complexity and scale of modern designs. Often, heuristics must trade off between turnaround time and quality of PPA outcomes. This paper presents a clustered placement methodology that improves both turnaround time and final-routed solution quality. Our PPA-aware clustering considers timing, power and logical hierarchy during netlist clustering, effectively reducing problem size and accelerating global placement runtime while improving post-route PPA metrics. Additionally, our machine learning (ML)-accelerated virtualized P&R methodology predicts the best cluster shapes (i.e., aspect ratios and utilizations) to use in P&R of the clustered netlist. With the open-source OpenROAD tool, our methods achieve up to 47% (average: 36%) global placement runtime improvement with similar half-perimeter wirelength (HPWL) and 90% (29%) improvement in post-route total negative slack (TNS). With the commercial Cadence Innovus tool, our methods achieve up to 3.92% (1%) improvement in power and 99% (49%) improvement in TNS. Andrew B. Kahng, Seokhyeong Kang, Sayak Kundu, Kyungjun Min, Seonghyeon Park, Bodhisatta Pramanik |
DAC | 1 |
| 2024 | Improvement of Mixed Track - Height Standard-Cell PlacementabstractIn sub-Snm nodes, track-height of standard cells must be aggressively scaled down while preserving design PPA. This requirement brings the challenge of placing a set of cells that have mixed track-heights, subject to the constraint that cells with the same height must be placed together in an “island” of cell rows. We apply integer linear programming (ILP) to solve the row assignment problem and improve the runtime of ILP with clustering, using a cost function that combines half-perimeter wirelength and displacement from a starting unconstrained placement. Considering the row assignment solution, we define fence-regions, which enable an existing place-and-route (P&R) tool to place the cells while considering the row-island constraints. Experimental results show that our proposed method can on average reduce final-routed wirelength by 8.5% and total power by 3.3 %, with worst negative slack and total negative slack reductions of 24.0% and 13.0%, compared with the previous state-of-the-art method [10]. Andrew B. Kahng, Seokhyeong Kang, Minhyuk Kweon |
DATE | 1 |
| 2024 | Scalable Flip-Flop Clustering Using Divide and Conquer For Capacitated K-MeansabstractMulti-bit flip-flop clustering is a well-studied optimization problem in physical design: carefully merging multiple single-bit flip-flops into a single multi-bit flip-flop can decrease the total power consumption in a clock distribution network due to lower clock power and routed clock wirelength. We propose a pointset decomposition heuristic that in conjunction with capacitated k-means [4] enables a scalable, divide-and-conquer flow for multi-bit flip-flop clustering. Our flow produces high-quality flip-flop clustering and placement solutions with respect to total power consumption, area, timing, and wirelength metrics evaluated after the post-routing optimization (PRO) stage of P&R. We test our flow on five designs of varying input size (0.5K to 64K clusterable single-bit flip-flops) implemented using the ASAP7 7nm research enablement [3] [9]. Empirical results show that our new flow is competitive with current state-of-the-art flows. Compared to MeanShift [2], we achieve 6.18% (resp. 1.90%) maximum (resp. average) reduction in total power consumption, along with improved total negative slack and wirelength. Compared to FlopTray [4], we achieve a 400 × speedup on larger designs such as VGA (17K single-bit flip-flops), but with an average 1.12% degradation in total power consumption. Andrew B. Kahng, Sayak Kundu, Shreyas Thumathy |
ACM Great Lakes Symposium on VLSI | 1 |
| 2024 | A Hybrid ECO Detailed Placement Flow for Improved Reduction of Dynamic IR DropabstractWith advanced semiconductor technology progressing well into sub-7nm scale, voltage drop has become an increasingly challenging issue. As a result, there has been extensive research focused on predicting and mitigating dynamic IR drops, leading to the development of IR drop engineering change order (ECO) flows – often integrated with modern commercial EDA tools. However, these tools encounter QoR limitations while mitigating IR drop. To address this, we propose a hybrid ECO detailed placement approach that is integrated with existing commercial EDA flows, to mitigate excessive peak current demands within power and ground rails. Our proposed hybrid approach effectively optimizes peak current levels within a specified “clip”– complementing and enhancing commercial EDA dynamic IR-driven ECO detailed placements. In particular, we: (i) order instances in a netlist in decreasing order of worst voltage drop; (ii) extract a clip around each instance; and (iii) solve an integer linear programming (ILP) problem to optimize instance placements. Our approach optimizes dynamic voltage drops (DVD) across ten designs by up to 15.3% compared to original conventional flows, with similar timing quality and 55.1% less runtime. Andrew B. Kahng, Bodhisatta Pramanik, Mingyu Woo |
ACM Great Lakes Symposium on VLSI | 1 |
| 2024 | Strengthening the Foundations for IC Physical Design and ML EDA ResearchabstractOver the past year, IEEE CEDA DATC has continued to improve the DATC Robust Design Flow (RDF) while also advancing open infrastructure for research, including machine learning for electronic design automation (ML EDA). The 2024 RDF release includes new standalone and integrated global placement and macro placement engines, as well as a CCS-based delay calculator. Advances in baselines and benchmarks include the addition of new benchmarks for macro placement and logic gate sizing, as well as further efforts to establish calibrations of both optimizations and analyses to aid assessments of research progress in EDA. Additional efforts to promote open and reproducible research include refined proxy research enablements and enhanced ML EDA infrastructure through the development and use of new formats, the release of datasets, and the development of Python APIs in OpenROAD. Vidya A. Chhabria, Vikram Gopalakrishnan, Andrew B. Kahng, Sayak Kundu, Zhiang Wang, Bing-Yue Wu, Dooseok Yoon |
ICCAD | 3 |
| 2024 | Placement Tomography-Based Routing Blockage Generation for DRV Hotspot MitigationabstractA fundamental goal in modern physical design is for the post-route layout to have a fixable number of remaining design rule violations (DRVs). We study how to apply routing blockages to a fixed placement solution, so as to "condition" the routing problem and minimize DRVs in the post-route outcome. Motivated by the widening turnaround time gap between early global routing (eGR) and detailed routing, we propose placement tomography (that uses multiple views of a placement from near-free eGR runs) as a new basis for generating layer-wise route blockages and mitigating post-route DRVs. Our framework includes (i) DRVNet, a machine learning model that predicts layer-wise DRV hotspots; (ii) BlkgComp, a learning-based model for assessing the relative effectiveness of two different routing blockages in mitigating DRVs in hotspots; and (iii) a reinforcement learning approach with BlkgComp to generate routing blockages for the hotspots predicted by DRVNet. Experimental studies confirm that our BlkgComp model achieves up to 73% accuracy and 0.53 Kendall rank on the testing dataset for open-source and commercial enablements. Our framework produces routing blockage solutions that reduce post-route DRVs by up to 88% compared to baseline commercial tool flows and up to 21% compared to a human expert baseline that was able to access detailed route outcomes. Andrew B. Kahng, Sayak Kundu, Dooseok Yoon |
ICCAD | 1 |
| 2024 | Panel Statement: EDA Needs at Advanced Technology NodesabstractImprovements in EDA technology become more urgent with the loss of traditional scaling levers and the growing complexity of leading-edge designs. The roadmap for device and cell architectures, power delivery, multi-tier integration, and multiphysics signoffs presents severe challenges to current optimization frameworks. In future advanced technologies, the quality of results, speed, and cost of EDA will be essential components of scaling. Each stage of the design flow must deliver predictable results for given inputs and targets. Optimization goals must stay correlated to downstream and final design outcomes. Andrew B. Kahng |
ISPD | 1 |
| 2024 | Solvers, Engines, Tools and Flows: The Next Wave for AI/ML in Physical DesignabstractIt has been six years since an ISPD-2018 invited talk on "Machine Learning Applications in Physical Design". Since then, despite considerable activity across both academia and industry, many R&D targets remain open. At the same time, there is now clearer understanding of where AI/ML can and cannot (yet) move the needle in physical design, as well as some of the difficult blockers and technical challenges that lie ahead. Some futures for AI/ML-boosted physical design are visible across solvers, engines, tools and flows - and in contexts that span generative AI, the modeling of "magic" handoffs at flow interstices, academic research infrastructure, and the culture of benchmarking and open-source EDA. Andrew B. Kahng |
ISPD | 1 |
| 2024 | OpenROAD and CircuitOps: Infrastructure for ML EDA Research and EducationabstractTraditional electronic design automation (EDA) techniques struggle to fulfill the stringent efficiency and quick turnaround demands of complex integrated systems. Machine learning (ML) strategies for EDA (“ML EDA”) are pivotal in transforming EDA to address these challenges. However, they encounter significant obstacles due to inadequate infrastructure, ranging from datasets to software interfaces. This paper demonstrates a software infrastructure for ML EDA built on two key technologies: (i) OpenROAD’s Python APIs, and (ii) NVIDIA’s CircuitOps, an EDA data representation format tailored for ML, facilitating ML EDA applications. The paper illustrates three ML EDA examples that utilize the established OpenROAD and CircuitOps infrastructure. Vidya A. Chhabria, Wenjing Jiang, Andrew B. Kahng, Rongjian Liang, Haoxing Ren, Sachin S. Sapatnekar, Bing-Yue Wu |
VTS | 3 |
| 2024 | K-SpecPart: Supervised Embedding Algorithms and Cut Overlay for Improved Hypergraph PartitioningabstractState-of-the-art hypergraph partitioners follow the multilevel paradigm that constructs multiple levels of progressively coarser hypergraphs that are used to drive cut refinement on each level of the hierarchy. Multilevel partitioners are subject to two limitations: 1) hypergraph coarsening processes rely on local neighborhood structure without fully considering the global structure of the hypergraph and 2) refinement heuristics risk entrapment in local minima. In this article, we describe K-SpecPart, a supervised spectral framework for multiway partitioning that directly tackles these two limitations. K-SpecPart relies on the computation of generalized eigenvectors and supervised dimensionality reduction techniques to generate vertex embeddings. These are computational primitives that are not only fast, but embeddings also capture global structural properties of the hypergraph that are not explicitly considered by existing partitioners. K-SpecPart then converts the vertex embeddings into multiple partitioning solutions. Unlike multilevel partitioners that only consider the best solution, K-SpecPart introduces the idea of “ensembling” multiple solutions via a cut-overlay clustering technique that often enables the use of computationally demanding partitioning methods such as integer linear programming (ILP). Using the output of a standard partitioner as a supervision hint, K-SpecPart effectively combines the strengths of established multilevel partitioning techniques with the benefits of spectral graph theory and other combinatorial algorithms. K-SpecPart significantly extends ideas and algorithms that first appeared in our previous work on the bipartitioner SpecPart (Bustany et al., ICCAD 2022). Our experiments demonstrate the effectiveness of K-SpecPart. For bipartitioning, K-SpecPart produces solutions with up to ~15% cutsize improvement over SpecPart. For multiway partitioning, K-SpecPart produces solutions with up to ~20% cutsize improvement for smaller$K$, and maintains ~2% improvement even when$K$is increased to 128, over leading partitioners hMETIS and KaHyPar. Ismail Bustany, Andrew B. Kahng, Ioannis Koutis, Bodhisatta Pramanik, Zhiang Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | PROBE3.0: A Systematic Framework for Design-Technology Pathfinding With Improved Design EnablementabstractWe propose a systematic framework to conduct design-technology pathfinding for power, performance, area, and cost (PPAC) in advanced nodes. Our goal is to provide a configurable, scalable generation of process design kit (PDK) and standard-cell library, spanning key scaling boosters (backside PDN and buried power rail), to explore PPAC across given technology and design parameters. We build on Cheng et al. (2022), which addressed only area and cost (AC), to include power and performance (PP) evaluations through automated generation of full design enablements. We also improve the use of artificial designs in the PPAC assessment of technology and design configurations. We generate more realistic artificial designs by applying a machine learning-based parameter tuning flow to Kim et al. (2022). We further employ clustering-based cell width-regularized placements at the core of routability assessment, enabling more realistic placement utilization and improved experimental efficiency. We evaluate PPAC across scaling boosters and artificial designs in a predictive technology node. Suhyeong Choi, Jinwook Jung, Andrew B. Kahng, Chul-Hong Park, Bodhisatta Pramanik, Dooseok Yoon |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | Hier-RTLMP: A Hierarchical Automatic Macro Placer for Large-Scale Complex IP BlocksabstractIn a typical RTL to GDSII flow, floorplanning or macro placement is a critical step in achieving decent quality of results (QoR). Moreover, in today’s physical synthesis flows (e.g., Synopsys Fusion Compiler or Cadence Genus iSpatial), a floorplan.def with macro and IO pin placements is typically needed as an input to the front-end physical synthesis. Recently, with the increasing complexity of IP blocks, and in particular with auto-generated RTL for machine learning (ML) accelerators, the number of macros in a single RTL block can easily run into the several hundreds. This makes the task of generating an automatic floorplan (.def) with IO pin and macro placements for front-end physical synthesis even more critical and challenging. The so-called peripheral approach of forcing macros to the periphery of the layout is no longer viable when the ratio of the sum of the macro perimeters to the floorplan perimeter is large, since this increases the required stacking depth of macros. In this article, we develop a novel multilevel physical planning approach that exploits the hierarchy and dataflow inherent in the design RTL, and describe its realization in a new hierarchical macro placer, Hier-RTLMP. Hier-RTLMP borrows from traditional approaches used in manual system-on-chip (SoC) floorplanning to create an automatic macro placement for use with large IP blocks containing very large numbers of macros. Empirical studies demonstrate substantial improvements over the previous RTL-MP macro placement approach (Kahng et al., 2022), and promising post-route improvements relative to a leading commercial place-and-route tool. Andrew B. Kahng, Ravi Varadarajan, Zhiang Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | A Machine Learning Approach to Improving Timing Consistency between Global Route and Detailed RouteabstractDue to the unavailability of routing information in design stages prior to detailed routing (DR), the tasks of timing prediction and optimization pose major challenges. Inaccurate timing prediction wastes design effort, hurts circuit performance, and may lead to design failure. This work focuses on timing prediction after clock tree synthesis and placement legalization, which is the earliest opportunity to time and optimize a “complete” netlist. The article first documents that having “oracle knowledge” of the final post-DR parasitics enables post-global routing (GR) optimization to produce improved final timing outcomes. To bridge the gap between GR-based parasitic and timing estimation and post-DR results during post-GR optimization , machine learning (ML)-based models are proposed, including the use of features for macro blockages for accurate predictions for designs with macros. Based on a set of experimental evaluations, it is demonstrated that these models show higher accuracy than GR-based timing estimation. When used during post-GR optimization, the ML-based models show demonstrable improvements in post-DR circuit performance. The methodology is applied to two different tool flows—OpenROAD and a commercial tool flow—and results on an open-source 45nm bulk and a commercial 12nm FinFET enablement show improvements in post-DR timing slack metrics without increasing congestion. The models are demonstrated to be generalizable to designs generated under different clock period constraints and are robust to training data with small levels of noise. Vidya A. Chhabria, Wenjing Jiang, Andrew B. Kahng, Sachin S. Sapatnekar |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2024 | An Open-Source ML-Based Full-Stack Optimization Framework for Machine Learning AcceleratorsabstractParameterizable machine learning (ML) accelerators are the product of recent breakthroughs in ML. To fully enable their design space exploration (DSE), we propose a physical-design-driven, learning-based prediction framework for hardware-accelerated deep neural network (DNN) and non-DNN ML algorithms. It adopts a unified approach that combines power, performance, and area (PPA) analysis with frontend performance simulation, thereby achieving a realistic estimation of both backend PPA and system metrics such as runtime and energy. In addition, our framework includes a fully automated DSE technique, which optimizes backend and system metrics through an automated search of architectural and backend parameters. Experimental studies show that our approach consistently predicts backend PPA and system metrics with an average 7% or less prediction error for the ASIC implementation of two deep learning accelerator platforms, VTA and VeriGOOD-ML, in both a commercial 12 nm process and a research-oriented 45 nm process. Hadi Esmaeilzadeh, Soroush Ghodrati, Andrew B. Kahng, Joon Kyung Kim, Sean Kinzer, Sayak Kundu, Rohan Mahapatra, Susmita Dey Manasi, Sachin S. Sapatnekar, Zhiang Wang, Ziqing Zeng |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2023 | An Open-Source Constraints-Driven General Partitioning Multi-Tool for VLSI Physical DesignabstractWith the increasing complexity of IC products, large-scale designs must be efficiently partitioned into multiple blocks, tiles, or devices for concurrent backend place-and-route (P&R) implementation. State-of-the-art partitioners focus on balanced min-cut without considering constraints such as timing or heterogeneity of resource types. They are thus increasingly unsuitable for current physical design requirements. We introduce TritonPart, the first open-source, constraints-driven partitioning tool for VLSI physical design. TritonPart employs efficient algorithms to handle constraints, including multi-dimensional balance, embedding, and timing constraints. Our experimental work affirms its benefits. For standard min-cut partitioning, TritonPart outperforms hMETIS [17], with improvements of up to ~20% on some benchmarks. For embedding-aware partitioning, TritonPart effectively leverages the embeddings generated by SpecPart [4] and improves upon it by ~2%. For timing-aware partitioning, TritonPart significantly reduces the number of cuts on timing-critical paths and prevents timing-noncritical paths from becoming critical (~21X, ~119X reduction relative to hMETIS and KaHyPar [31], respectively). Ismail Bustany, Grigor Gasparyan, Andrew B. Kahng, Ioannis Koutis, Bodhisatta Pramanik, Zhiang Wang |
ICCAD | 3 |
| 2023 | Invited Paper: The Inevitability of AI Infusion Into Design Closure and SignoffabstractSoC design teams embrace new technologies and methodologies that bring clear value. Given this, future infusion of AI into design closure and signoff is inevitable. Predictive AI models help focus the application of last-mile incremental optimizations (sizing, placement and routing) to achieve timing and noise closure; successful examples range from routing-free crosstalk prediction to timing/power evaluation in early RTL development. Design closure becomes more efficient when “imperfect but fast” ML inferencing is used to filter out potential violations, which can then be passed to golden analysis tools. Learning methods also improve the design process in many ways, ranging from smarter PVT corner selection to predicting the CPU and memory usage of signoff tools. At a higher level, AI will help design teams learn to avoid design trajectories that lead to time-consuming closure and signoff iterations. This talk will provide a broad overview of directions in which AI will inevitably improve the cost and efficiency of signoff in the coming years. Jiang Hu 0001, Andrew B. Kahng |
ICCAD | 2 |
| 2023 | Invited Paper: IEEE CEDA DATC Emerging Foundations in IC Physical Design and MLCAD ResearchabstractRecent activities of the IEEE CEDA DATC strengthen the DATC Robust Design Flow (RDF) and broadly support research on machine learning for CAD/EDA (MLCAD). The RDF-2023 version of the RDF adds standalone and integrated netlist partitioners, a detailed placement optimizer, dynamic power analysis, and enablement of new directions (design-technology co-optimization and 3D layout). Advancement of benchmarking practices and strong baselines has continued - e.g., the MacroPlacement effort introduced in RDF-2022 now has new benchmarks, integration of the AutoDMP macro placer, and baseline solutions generated by Simulated Annealing and human experts. Other DATC efforts have focused on proxies and other elements of MLCAD research enablement. These include real and synthetic benchmarks tailored for IR drop analysis, a calibration methodology for research PDKs, and artificial netlist generation for data augmentation and design space coverage of netlists used in model training. We conclude with directions for future DATC efforts. Jinwook Jung, Andrew B. Kahng, Sayak Kundu, Zhiang Wang, Dooseok Yoon |
ICCAD | 2 |
| 2023 | Assessment of Reinforcement Learning for Macro PlacementabstractWe provide open, transparent implementation and assessment of Google Brain's deep reinforcement learning approach to macro placement (Nature) and its Circuit Training (CT) implementation in GitHub. We implement in open-source key "blackbox" elements of CT, and clarify discrepancies between CT and Nature. New testcases on open enablements are developed and released. We assess CT alongside multiple alternative macro placers, with all evaluation flows and related scripts public in GitHub. Our experiments also encompass academic mixed-size placement benchmarks, as well as ablation and stability studies. We comment on the impact of Nature and CT, as well as directions for future research. Chung-Kuan Cheng, Andrew B. Kahng, Sayak Kundu, Yucheng Wang 0016, Zhiang Wang |
ISPD | 2 |
| 2023 | DAGSizer: A Directed Graph Convolutional Network Approach to Discrete Gate Sizing of VLSI GraphsabstractThe objective of a leakage recovery step is to make use of positive slack and reduce power by performing appropriate standard-cell swaps such as threshold-voltage ( V th ) or channel-length reassignments. The resulting engineering change order netlist needs to be timing clean. Because this recovery step is performed several times in a physical design flow and involves long runtimes and high tool-license usage, previous works have proposed graph neural network–based frameworks that restrict feature aggregation to three-hop neighborhoods and do not fully consider the directed nature of netlist graphs. As a result, the intermediate node embeddings do not capture the complete structure of the timing graph. In this article, we propose DAGSizer , a framework that exploits the directed acyclic nature of timing graphs to predict cell reassignments in the discrete gate sizing task. Our DAGSizer (Sizer for DAGs) framework is based on a node ordering-aware recurrent message-passing scheme for generating the latent node embeddings. The generated node embeddings absorb the complete information from the fanin cone (predecessors) of the node. To capture the fanout information into the node embeddings, we enable a bidirectional message-passing mechanism. The concatenated latent node embeddings from the forward and reverse graphs are then translated to nodewise delta-delay predictions using a teacher sampling mechanism. With eight possible cell-assignments, the experimental results demonstrate that our model can accurately estimate design-level leakage recovery with an absolute relative error ε model under 5.4%. As compared to our previous work, GRA-LPO, we also demonstrate a significant improvement in the model mean squared error. Chung-Kuan Cheng, Chester Holtz, Andrew B. Kahng, Bill Lin 0001, Uday Mallappa |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2022 | AI/ML, Optimization and EDA in the TILOS AI Research InstituteabstractOptimization means finding the best possible solution to a given problem -- but challenges of scale and complexity keep many real-world optimization needs beyond reach. The Institute for Learning-enabled Optimization at Scale (TILOS) is a new National AI Research Institute led by UC San Diego in partnership with MIT, National University, Penn, UT Austin and Yale. TILOS is sponsored by the U.S. National Science Foundation, with partial support from Intel Corporation. The TILOS mission: make impossible optimizations possible, at scale and in practice. This talk introduces TILOS, its research agenda, and how it aims to establish a "national nexus" of AI and machine learning, optimization, and use domains that include integrated-circuit design and design automation. The talk will point out directions along which the interplay of learning and optimization can boost the scaling and quality of EDA outcomes. Andrew B. Kahng |
ACM Great Lakes Symposium on VLSI | 1 |
| 2022 | SpecPart: A Supervised Spectral Framework for Hypergraph Partitioning Solution ImprovementabstractState-of-the-art hypergraph partitioners follow the multilevel paradigm that constructs multiple levels of progressively coarser hypergraphs that are used to drive cut refinements on each level of the hierarchy. Multilevel partitioners are subject to two limitations: (i) Hypergraph coarsening processes rely on local neighborhood structure without fully considering the global structure of the hypergraph. (ii) Refinement heuristics can stagnate on local minima. In this paper, we describe SpecPart, the first supervised spectral framework that directly tackles these two limitations. SpecPart solves a generalized eigenvalue problem that captures the balanced partitioning objective and global hypergraph structure in a low-dimensional vertex embedding while leveraging initial high-quality solutions from multilevel partitioners as hints. SpecPart further constructs a family of trees from the vertex embedding and partitions them with a tree-sweeping algorithm. Then, a novel overlay of multiple tree-based partitioning solutions, followed by lifting to a coarsened hypergraph, where an ILP partitioning instance is solved to alleviate local stagnation. We have validated SpecPart on multiple sets of benchmarks. Experimental results show that for some benchmarks, our SpecPart can substantially improve the cutsize by more than 50% with respect to the best published solutions obtained with leading partitioners hMETIS and KaHyPar. Ismail Bustany, Andrew B. Kahng, Ioannis Koutis, Bodhisatta Pramanik, Zhiang Wang |
ICCAD | 2 |
| 2022 | IEEE CEDA DATC: Expanding Research Foundations for IC Physical Design and ML-Enabled EDAabstractThis paper describes new elements in the RDF-2022 release of the DATC Robust Design Flow, along with other activities of the IEEE CEDA DATC. The RosettaStone initiated with RDF-2021 has been augmented to include 35 benchmarks and four open-source technologies (ASAP7, NanGate45 and SkyWater130HS/HD), plus timing-sensible versions created using path-cutting. The Hier-RTLMP macro placer is now part of DATC RDF, enabling macro placement for large modern designs with hundreds of macros. To establish a clear baseline for macro placers, new open-source benchmark suites on open PDKs, with corresponding flows for fully reproducible results, are provided. METRICS2.1 infrastructure in OpenROAD and OpenROAD-flow-scripts now uses native JSON metrics reporting, which is more robust and general than the previous Python script-based method. Calibrations on open enablements have also seen notable updates in the RDF. Finally, we also describe an approach to establishing a generic, cloud-native large-scale design of experiments for ML-enabled EDA. Our paper closes with future research directions related to DATC's efforts. Jinwook Jung, Andrew B. Kahng, Ravi Varadarajan, Zhiang Wang |
ICCAD | 2 |
| 2022 | A Mixed Open-Source and Proprietary EDA Commons for Education and PrototypingabstractIn recent years, several open-source projects have shown potential to serve a future technology commons for EDA and design prototyping. This paper examines how open-source and proprietary EDA technologies will inevitably take on complementary roles within a future technology commons. Proprietary EDA technologies offer numerous benefits that will endure, including (i) exceptional technology and engineering; (ii) ever-increasing importance in design-based equivalent scaling and the overall semiconductor value chain; and (iii) well-established commercial and partner relationships. On the other hand, proprietary EDA technologies face challenges that will also endure, including (i) inability to pursue directions such as massive leverage of cloud compute, extreme reduction of turnaround times, or "free tools"; and (ii) difficulty in evolving and addressing new applications and markets. By contrast, open-source EDA technologies offer benefits that include (i) the capability to serve as a friction-free, democratized platform for education and future workforce development (i.e., as a platform for EDA research, and as a means of teaching / training both designers and EDA developers with public code); and (ii) addressing the needs of underserved, non-enterprise account markets (e.g., older nodes, research flows, cost-sensitive IoT, new devices and integrations, system-design-technology pathfinding). This said, open-source will always face challenges such as sustainability, governance, and how to achieve critical mass and critical quality. The paper will conclude with key directions and synergies for open-source and proprietary EDA within an EDA Commons for education and prototyping. Andrew B. Kahng |
ICCAD | 1 |
| 2022 | Leveling Up: A Trajectory of OpenROAD, TILOS and BeyondabstractSince June 2018, the OpenROAD project has developed an open-source, RTL-to-GDS EDA system within the DARPA IDEA program. The tool achieves no-human-in-loop generation of design-rule clean layout in 24 hours. This enables system innovation and design space exploration, while also democratizing hardware design by lowering barriers of cost, expertise and risk. Since November 2021, The Institute for Learning-enabled Optimization at Scale (TILOS), an NSF AI institute for advances in optimization partially supported by Intel, has begun its work toward a "new nexus'' of AI, optimization, and the leading edge of practice for use domains that include IC design. This paper traces a trajectory of "leveling up'' in the research enablement for IC physical design automation and EDA in general. This trajectory has OpenROAD and TILOS as waypoints, and advances themes of openness, infrastructure, and culture change. Andrew B. Kahng |
ISPD | 1 |
| 2022 | RTL-MP: Toward Practical, Human-Quality Chip Planning and Macro PlacementabstractIn a typical RTL-to-GDSII flow, floorplanning plays an essential role in achieving decent quality of results (QoR). A good floorplan typically requires interaction between the frontend designer, who is responsible for the functionality of the RTL, and the backend physical design engineer. The increasing complexity of macro-dominated designs (especially machine learning accelerators with autogenerated RTL) has made the floorplanning task even more challenging and time-consuming. In this paper, we propose RTL-MP, a novel macro placer which utilizes RTL information and tries to "mimic" the interaction between the frontend RTL designer and the backend physical design engineer to produce human-quality floorplans. By exploiting the logical hierarchy and processing logical modules based on connection signatures, RTL-MP can capture the dataflow inherent in the RTL and use the dataflow information to guide macro placement. We also apply autotuning to optimize hyperparameter settings based on input designs. We have built RTL-MP based on OpenROAD infrastructure and applied RTL-MP to a set of industrial designs. RTL-MP outperforms state-of-the-art commercial macro placers and achieves QoR similar to that of handcrafted floorplans. Andrew B. Kahng, Ravi Varadarajan, Zhiang Wang |
ISPD | 1 |
| 2022 | PROBE2.0: A Systematic Framework for Routability Assessment From Technology to Design in Advanced NodesabstractIn advanced nodes, scaling of critical dimension and pitch has not progressed at historical Moore’s Law rates. Thus,scaling boostersare explored to improve achievable power, performance, area, and cost (PPAC) in new technologies. However, scaling boosters increase complexity of standard-cell architectures, power delivery, design rules, and other aspects of the design enablement, and may not result in design-level benefits. Therefore, design-technology co-optimization (DTCO) methodologies are required to evaluate design-level benefits of scaling boosters. The key challenge for DTCO is that large engineering efforts and long timelines are needed to develop design enablements (e.g., cell libraries) and perform implementation studies in order to assess technology options. We describe a new framework that can systematically evaluate a measure of intrinsic routability,$K_{\mathrm{ th}}$, across both technology and design choices. We focus on routability since it is a critical factor in the scaling of area and cost. Our framework includes realistic standard-cell libraries that are automatically generated using satisfiability modulo theory (SMT) methods, and a new pin shape selection method. Routability assessments are based on the PROBE approach and an improved construction of underlying netlist topologies. Our experimental studies demonstrate the assessment of routability impacts for advanced-node technology and design options. We demonstrate learning-based$K_{\mathrm{ th}}$prediction to reduce runtime, disk space and commercial tool licenses needed to implement our framework. Our work enables faster and more comprehensive evaluation of technology options early in the technology development process. Chung-Kuan Cheng, Andrew B. Kahng, Hayoung Kim, Daeyeal Lee, Dongwon Park, Mingyu Woo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | In-Route Pin Access-Driven Placement Refinement for Improved Detailed Routing ConvergenceabstractPin access is increasingly important in advanced nodes. Neighboring or cell-boundary pins can have degraded pin accessibility, causing design rule violations (DRCs) during routing, which are runtime expensive to resolve. Conventional physical design tool flow uses pessimistic and/or inaccurate understanding of pin access during the placement stage and keeps the location of cells fixed during routing. This can leave pin access issues unsolvable and block further routing solution improvement. The timeliness of our present work is confirmed by the recent ICCAD-2020 CAD Contest, Problem B formulation from Synopsys, Inc. (Hu and Yang, 2020). The organizers give a succinct motivation for what we study—to eliminate preserved margins and misalignment issues from conventional placement models. In this work, we develop anin-route, pin access-driven local placement refinement. Experiments across industry designs in a wide range of advanced technology nodes confirm that our optimization can significantly improve routing convergence (i.e., subsequent detailed routing runtime and initial detailed routing DRCs). Our optimization can reduce congestion and wirelength without timing degradation. Andrew B. Kahng, Wen-Hao Liu 0001, Bangqi Xu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | TritonRoute-WXL: The Open-Source Router With Integrated DRC EngineabstractRouting is a crucial stage in a modern design automation tool flow for advanced technology nodes. Works in the recent open literature tend to divide routing into separate global routing (GR) and detailed routing (DR) steps without addressing the correlation issues (e.g., local nets) between these two steps. In this work, we present TritonRoute-WXL (TR-WXL), a unified global-detailed router capable of delivering design rule check-clean routing solutions in commercial sub-16-nm technologies. The major contributions of TR-WXL include an end-to-end routing framework that closely connects GR and DR, and an improved DR flow. With a code release under a permissive open-source license, TR-WXL achieves unparalleled solution quality in terms of DR and global-detailed routing (GDR) as compared to known best solutions from all published academic routers. Andrew B. Kahng, Lutong Wang, Bangqi Xu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | DATC RDF-2021: Design Flow and Beyond ICCAD Special Session PaperabstractThis paper describes the latest release of the DATC Robust Design Flow (RDF), RDF-2021, which has several key additions to expand its horizons. The Chisel/FIRRTL compiler is now part of DATC RDF, enabling support of recent hardware generator designs written in Chisel. Logic locking through RTL obfuscation, an updated ABC synthesis flow, and DFT support are other notable updates to the RDF. A Bookshelf-LEF/DEF converter powered by OpenDB is also added into DATC RDF's inventory as an enabler of robust benchmark conversion. We also describe efforts toward open metrics standards and datasets for machine learning (ML) applications and smart tuning of the design flow, as well as expansion of public analysis calibration data. Our paper closes with future research directions related to DATC's efforts. Jianli Chen, Iris Hui-Ru Jiang, Jinwook Jung, Andrew B. Kahng, Seungwon Kim, Victor N. Kravets, Yih-Lang Li, Ravi Varadarajan, Mingyu Woo |
ICCAD | 4 |
| 2021 | VeriGOOD-ML: An Open-Source Flow for Automated ML Hardware SynthesisabstractThis paper introduces VeriGOOD-ML, an automated methodology for generating Verilog with no human in the loop, starting from a high-level description of a machine learning (ML) algorithm in a standard format such as ONNX. The Verilog RTL is then translated through a back-end design flow to GDSII, driven by a design planning approach that is well tailored to the macro-intensive nature of ML platforms. VeriGOOD-ML uses three approaches to build ML hardware: the TABLA platform uses a dataflow architecture that is well suited to non-DNN ML algorithms; the GeneSys platform, with a systolic array and a SIMD array, is optimized for implementing DNNs; and the Axiline approach synthesizes small ML algorithms by hardcoding the structure of the algorithm into hardware, thus trading off flexibility for performance and power. The overall approach explores the design space of platform configurations and Pareto-optimal-PPA back-end implementations to yield designs that represent different tradeoffs at the algorithmic level between area, power, performance, and execution time. The overall methodology, from architecture to back-end design to hardware implementation, is described in this paper, and the results of VeriGOOD-ML are demonstrated on a set of ML benchmarks. Hadi Esmaeilzadeh, Soroush Ghodrati, Jie Gu 0003, Andrew B. Kahng, Joon Kyung Kim, Sean Kinzer, Rohan Mahapatra, Susmita Dey Manasi, Edwin Mascarenhas, Sachin S. Sapatnekar, Ravi Varadarajan, Zhiang Wang, Hanyang Xu 0002, Brahmendra Reddy Yatham, Ziqing Zeng |
ICCAD | 5 |
| 2021 | METRICS2.1 and Flow Tuning in the IEEE CEDA Robust Design Flow and OpenROAD ICCAD Special Session PaperabstractIn today's RTL-to-GDS flow domain, there is a lack of standards for reporting of design and tool metrics. Moreover, each tool or engine has its own set of parameters that can change outcomes and trade off PPA and other metrics. Thus, the study and optimization of impacts of parameter settings across the entire RTL- to-GDS tool chain has been largely ad hoc. In this paper, we first describe METRICS2.1, a proposed standard for RTL-to-GDS design tool and flow metrics. We then describe how data collected using a METRICS2.1 realization can be analyzed to give insight into flow tuning and fields of use for PPA optimization. Last, we discuss hyperparameter autotuning in the RTL-to-GDS flow. We present AutoTuner, which uses derivative-free optimization to handle challenges of non-differentiability and many local minima. An open repository based on METRICS2.1 has been established for sharing of reproducible, standardized metrics data, along with example implemented applications, to support academic and industrial research on machine learning for tool/flow tuning. Jinwook Jung, Andrew B. Kahng, Seungwon Kim, Ravi Varadarajan |
ICCAD | 2 |
| 2021 | CoRe-ECO: Concurrent Refinement of Detailed Place-and-Route for an Efficient ECO AutomationabstractWith the relentless scaling of technology nodes, physical design engineers encounter non-trivial challenges caused by rapidly increasing design complexity, particularly in the routing stage. Back-end designers must manually stitch/modify all of the design rule violations (DRVs) that remain after automatic place-and-route (P&R), during the implementation of engineering change orders (ECOs). In this paper, we propose CoRe-ECO, a concurrent refinement framework for efficient automation of the ECO process. Our framework efficiently resolves pin accessibility-induced DRVs by simultaneously performing detailed placement, detailed routing, and cell replacement. In addition to perturbation-minimized solutions, our proposed SMT-based optimization framework also suggests the adoption of alternative master cells to better achieve DRV-clean layouts. We demonstrate that our framework successfully resolves from 33.3% to 100.0% (58.6% on average) of remaining DRVs on M1-M3 layers, across a range of benchmark circuits with various cell architectures, while also providing average total wirelength reduction of 0.003%. Chung-Kuan Cheng, Andrew B. Kahng, Ilgweon Kang, Daeyeal Lee, Bill Lin 0001, Dongwon Park, Mingyu Woo |
ICCD | 2 |
| 2021 | Advancing PlacementabstractPlacement is central to IC physical design: it determines spatial embedding, and hence parasitics and performance. From coarse-to fine-grain, placement is conjointly optimized with logic, performance, clock and power distribution, routability and manufacturability. This paper gives some personal thoughts on futures for placement research in IC physical design. Revisiting placement as optimization prompts a new look at placement requirements, optimization quality, and scalability with resources. Placement must also evolve to meet a growing need for co-optimizations and for co-operation with other design steps. "New" challenges will naturally arise from scaling, both at the end of the 2D scaling roadmap and in the context of future 2.5D/3D/4D integrations. And, the nexus of machine learning and placement optimization will continue to be an area of intense focus for research and practice. In general, placement research is likely to see more flow-scale optimization contexts, open source, benchmarking of progress toward optimality, and attention to translations into real-world practice. Andrew B. Kahng |
ISPD | 1 |
| 2021 | TritonRoute: The Open-Source Detailed RouterabstractDetailed routing is a dead-or-alive critical element in design automation tooling for advanced node enablement. However, very few works address detailed routing in the recent open literature, particularly in the context of modern industrial designs and a complete, end-to-end flow. The ISPD-2018 Initial Detailed Routing Contest addressed this gap for modern industrial designs, using a realistic design rules set. In this work, we present TritonRoute, a detailed router capable of delivering a DRC-clean routing solution. The key contributions of TritonRoute include an in-memory router database, along with an end-to-end detailed routing scheme that is capable of comprehending connectivity and design rule constraints, with every key detail revealed by a code release under a permissive open-source license. We evaluate our router using the official ISPD-2018 benchmark suite and show that TritonRoute achieves an unprecedented solution quality-improved wirelength and via count, and an extremely low level of design rule violations (DRCs). Compared to the known best detailed routing solutions from all published academic detailed routers, TritonRoute improves wirelength by up to 0.8% (avg. 0.4%), via count by up to 16.1% (avg. 9.3%), and DRCs by up to 100% (avg. 92.0%). Andrew B. Kahng, Lutong Wang, Bangqi Xu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | Enhanced Power Delivery Pathfinding for Emerging 3-D Integration TechnologyabstractIn advanced technology nodes, emerging 3-D integration technology is a promising “More Than Moore” lever for continued scaling of system capability and value. In the 3-D integrated circuit (3-D IC) implementation, the power delivery network (PDN) is crucial to meeting design specifications. However, determining the optimal PDN design is nontrivial. On the one hand, to meet the voltage (IR) drop requirement, a denser power mesh is desired. On the other hand, to meet the timing requirement, more routing resource is needed for signal routing. Moreover, additional competition between signal routing and power routing is caused by intertier vertical interconnects in 3-D IC. In this article, we propose a power delivery pathfinding methodology for emerging 3-D integration, which seeks to identify a “near-optimal” (or, very high quality) PDN for a given BEOL stack, vertical interconnection, and PDN specification. Compared with previous works, our methodology can explore richer solution spaces as it supports different PDN layer combinations and PDN layer configurations. We develop models for routability and worst IR drop to help reduce iterations between PDN design and circuit design in 3-D IC implementation. We present validations and demonstrate improvement in IR drop and routability with real design blocks in 28- and 14-nm foundry technology nodes. Andrew B. Kahng, Seokhyeong Kang, Seungwon Kim, Bangqi Xu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2020 | Template-based PDN Synthesis in Floorplan and Placement Using Classifier and CNN TechniquesabstractDesigning an optimal power delivery network (PDN) is a time-intensive task that involves many iterations. This paper proposes a methodology that employs a library of predesigned, stitchable templates, and uses machine learning (ML) to rapidly build a PDN with region-wise uniform pitches based on these templates. Our methodology is applicable at both the floorplan and placement stages of physical implementation. (i) At the floorplan stage, we synthesize an optimized PDN based on early estimates of current and congestion, using a simple multilayer perceptron classifier. (ii) At the placement stage, we incrementally optimize an existing PDN based on more detailed congestion and current distributions, using a convolution neural network. At each stage, the neural network builds a safe-by-construction PDN that meets IR drop and electromigration (EM) specifications. On average, the optimization of the PDN brings an extra 3% of routing resources, which corresponds to a thousands of routing tracks in congestion-critical regions, when compared to a globally uniform PDN, while staying within the IR drop and EM limits. Vidya A. Chhabria, Andrew B. Kahng, Uday Mallappa, Sachin S. Sapatnekar, Bangqi Xu |
ASP-DAC | 2 |
| 2020 | The Tao of PAO: Anatomy of a Pin Access Oracle for Detailed RoutingabstractPin accessibility has been widely studied, particularly in recent works that span detailed placement optimization, standard cell layout optimization and new design rule-aware access model. However, to our knowledge, no previous work has described a full solution for pin access analysis, with validations on real detailed routing benchmarks. This paper presents a complete, robust, scalable and design rule-aware dynamic programming-based pin access analysis framework that is capable of both standard cell-based and instance-based pin access analysis. Integration into the open-source TritonRoute router results in superior solution quality compared to previous best-known results for the official ISPD-2018 benchmark suite. Andrew B. Kahng, Lutong Wang, Bangqi Xu |
DAC | 1 |
| 2020 | DATC RDF-2020: Strengthening the Foundation for Academic Research in IC Physical DesignabstractWe describe the RDF-2020 release of the IEEE CEDA DATC Robust Design Flow (RDF). RDF-2020 extends the previous four years of DATC efforts to (i) preserve and integrate leading research codes, including from past academic contests, and (ii) provide a foundation and backplane for academic research in the RTL-to-GDS IC implementation arena. Implementation and analysis flows have been enhanced by the addition of steps including multi-bit flip-flop clustering, parasitic extraction and antenna checking, as well as a recent contest-winning global router. RDF-2020 also opens a new "Calibrations" direction to support academic research on key analyses such as extraction and timing. An open-source physical design database with Tcl/Python/C++ APIs, a flow integration into a single scriptable application, and support for the newly-opened SKY130 manufacturable PDK, are also new this year. Our paper closes with a discussion of potential future directions for the RDF effort. Jianli Chen, Iris Hui-Ru Jiang, Jinwook Jung, Andrew B. Kahng, Victor N. Kravets, Yih-Lang Li, Shih-Ting Lin, Mingyu Woo |
ICCAD | 4 |
| 2020 | Open-Source EDA: If We Build It, Who Will Come?abstractThe VLSI technology and scaling roadmap has always included Process technology (wrapped as “PDK”), VLSI designs themselves (“System Drivers”), and EDA technology (“Design Technology”). Today, we see an open-source foundry PDK, and we see a vibrant open-source hardware design ecosystem. But what about open-source EDA? The development of open-source EDA technology cannot be separated from the question, “If we build it, who will come?” Today's talk will try to provide some thoughts on this question. What is “it”? Who is “we”? Who is “who”? And in what ways will the “who” come to interact with open-source EDA? Andrew B. Kahng |
VLSI-SOC | 1 |
| 2020 | On the superiority of modularity-based clustering for determining placement-relevant clusters
Mateus Fogaça, Andrew B. Kahng, Eder Monteiro, Ricardo Augusto da Luz Reis, Lutong Wang, Mingyu Woo |
Integr. | 2 |
| 2020 | Cross-Layer Co-Optimization of Network Design and Chiplet Placement in 2.5-D Systemsabstract2.5-D integration technology is gaining attention and popularity in manycore computing system design. 2.5-D systems integrate homogeneous or heterogeneous chiplets in a flexible and cost-effective way. The design choices of 2.5-D systems impact overall system performance, manufacturing cost, and thermal feasibility. This article proposes a cross-layer co-optimization methodology for 2.5-D systems. We jointly optimize the network topology and chiplet placement across logical, physical, and circuit layers to improve system performance, reduce manufacturing cost, and lower operating temperature, while ensuring thermal safety and routability. We also propose a novel gas-station link, which enables pipelined interchiplet links in passive interposers. Our cross-layer methodology achieves better performance-cost tradeoffs of 2.5-D systems and yields better solutions in optimizing interchiplet network and 2.5-D system designs than prior methods. Compared to single-chip systems, 2.5-D systems designed using our new approach achieve 88% higher performance at the same manufacturing cost, or 29% lower cost with the same performance. Compared to the closest state-of-the-art, our new approach achieves 40%-68% (49% on average) iso-cost performance improvement and 30%-38% (32% on average) iso-performance cost reduction. Ayse K. Coskun, Furkan Eris, Ajay Joshi, Andrew B. Kahng, Yenai Ma, Aditya Narayan, Vaishnav Srinivas |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Heuristic Methods for Fine-Grain Exploitation of FDSOIabstractFully depleted silicon on insulator (FDSOI) is attractive for its low cost and low power; the mixed-Vt and body-bias levers that it affords expand the performance-power solution space. However, in FDSOI, different-Vt (i.e., low Vt and regular Vt) devices must be isolated from each other, which makes realization of fine-grained mixed-Vt/body-biasing in layout extremely challenging. In this article, we study heuristic methods aimed at exploitation of fine-grained mixed-Vt in FDSOI implementation. We propose a novel speed domain partitioning (SDP) problem formulation that comprehends the spatial contiguity restrictions arising from flip-well structure of low Vt regions in popular 28-nm commercial FDSOI offerings. We explore a wide space of implementation flows that include an integer linear programming (ILP)-based approach, and a heuristic (sensitivity-based) optimization. Our experimental studies have been performed across multiple commercial enablements. We observe that outcomes are library- and design-dependent. For implementations using generic library options, up to 20% speed improvement with 54% low Vt region is seen for one out of four testcases studied. For implementations using “rich” library options, up to 7% speed improvement with 26% low Vt region is achieved. We provide a discussion that summarizes root-cause, intrinsic difficulties of fine-grained exploitation of mixed-Vt. Finally, we suggest a “decision tree” to help assess a design's amenability to fine-grained mixed-Vt implementation, and to help guide design flow selection for better design QoR. Hamed Fatemi, Andrew B. Kahng, Hyein Lee 0001, José Pineda de Gyvez |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Optimal Generalized H-Tree Topology and Buffering for High-Performance and Low-Power Clock DistributionabstractClock power, skew and maximum latency are three key metrics for clock distribution in low-power and high-performance designs. An H-tree offers minimum clock skew and good robustness against variations, but at the cost of large wirelength and clock power. On the other hand, a “fishbone” clock network with spine-ribs structures has smaller wirelength, latency and clock power, but larger skew, as compared to an H-tree. No previous work enables systematic exploration of the regime between H-tree and spine to achieve an optimal tradeoff among clock power, skew, and latency. In this paper, we study the concept of a generalized H-tree (GH-tree)-a topologically balanced tree with an arbitrary sequence of branching factors-and propose a dynamic programming-based method to determine optimal clock power, skew, and latency, in the space of GH-tree solutions. Our method co-optimizes clock tree topology and buffering along branches according to fitted electrical models. We further propose a balanced K-means clustering and a linear programming (LP)-guided buffer placement approach to embed the GH-tree with respect to a given sink placement. We validate our solutions in commercial clock tree synthesis (CTS) tool flows, in a commercial foundry's 28LP technology. The results show up to 30% clock power reduction while achieving similar skew and maximum latency as CTS solutions from recent versions of leading commercial place-and-route tools. Our proposed approach also achieves up to 56% clock power reduction while achieving similar skew and maximum latency as compared to CTS solutions from a state-of-the-art academic tool. Kwangsoo Han, Andrew B. Kahng, Jiajia Li 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Learning-based prediction of package power delivery network qualityabstractPower Delivery Network (PDN) is a critical component in modern System-on-Chip (SoC) designs. With the rapid development in applications, the quality of PDN, especially Package (PKG) PDN, determines whether a sufficient amount of power can be delivered to critical computing blocks. In conventional PKG design, PDN design typically takes multiple weeks including many manual iterations for optimization. Also, there is a large discrepancy between (i) quick simulation tools used for quick PDN quality assessment during the design phase, and (ii) the golden extraction tool used for signoff. This discrepancy may introduce more iterations. In this work, we propose a learning-based methodology to perform PKG PDN quality assessment both before layout (when only bump/ball maps, but no package routing, are available) and after layout (when routing is completed but no signoff analysis has been launched). Our contributions include (i) identification of important parameters to estimate the achievable PKG PDN quality in terms of bump inductance; (ii) the avoidance of unnecessary manual trial and error overheads in PKG PDN design; and (iii) more accurate design-phase PKG PDN quality assessment. We validate accuracy of our predictive models on PKG designs from industry. Experimental results show that, across a testbed of 17 industry PKG designs, we can predict bump inductance with an average absolute percentage error of 21.2% or less, given only pinmap and technology information. We improve prediction accuracy to achieve an average absolute percentage error of 17.5% or less when layout information is considered. Andrew B. Kahng, Joseph Li, Abinash Roy, Vaishnav Srinivas, Bangqi Xu |
ASP-DAC | 2 |
| 2019 | Finding placement-relevant clusters with fast modularity-based clusteringabstractIn advanced technology nodes, IC implementation faces increasing design complexity as well as ever-more demanding design schedule requirements. This raises the need for new decomposition approaches that can help reduce problem complexity, in conjunction with new predictive methodologies that can help avoid bottlenecks and loops in the physical implementation flow. Notably, with modern design methodologies it would be very valuable to better predict final placement of the gate-level netlist: this would enable more accurate early assessment of performance, congestion and floorplan viability in the SOC floorplanning/RTL planning stages of design. In this work, we study a new criterion for the classic challenge of VLSI netlist clustering: how well netlist clusters "stay together" through final implementation. We propose use of several evaluators of this criterion. We also explore the use of modularity-driven clustering to identify natural clusters in a given graph without the tuning of parameters and size balance constraints typically required by VLSI CAD partitioning methods. We find that the netlist hypergraph-to-graph mapping can significantly affect quality of results, and we experimentally identify an effective recipe for weighting that also comprehends topological proximity to I/Os. Further, we empirically demonstrate that modularity-based clustering achieves better correlation to actual netlist placements than traditional VLSI CAD methods (our method is also 4X faster than use of hMetis for our largest testcases). Finally, we show a potential flow with fast "blob placement" of clusters to evaluate netlist and floorplan viability in early design stages; this flow can predict gate-level placement of 370K cells in 200 seconds on a single core. Mateus Fogaça, Andrew B. Kahng, Ricardo Augusto da Luz Reis, Lutong Wang |
ASP-DAC | 2 |
| 2019 | Diffusion break-aware leakage power optimization and detailed placement in sub-10nm VLSIabstractA diffusion break (DB) isolates two neighboring devices in a standard cell-based design and has a stress effect on delay and leakage power. In foundry sub-10nm design enablements, device performance is changed according to the type of DB - single diffusion break (SDB) or double diffusion break (DDB) - that is used in the library cell layout. Crucially, local layout effect (LLE) can substantially affect device performance and leakage. Our present work focuses on the 2nd DB effect, a type of LLE in which distance to the second-closest DB (i.e., a distance that depends on the placement of a given cell's neighboring cell) also impacts performance of a given device. In this work, we implement a 2nd DB-aware timing and leakage analysis flow, and show how a lack of 2nd DB awareness can misguide current optimization in place-and-route stages. We then develop 2nd DB-aware leakage optimization and detailed placement heuristics. Experimental results in a scaled foundry 14nm technology indicate that our 2nd DB-aware analysis and optimization flow achieves, on average, 80% recovery of the leakage increment that is induced by the 2nd DB effect, without changing design performance. Sun ik Heo, Andrew B. Kahng, Lutong Wang |
ASP-DAC | 2 |
| 2019 | Toward an Open-Source Digital Flow: First Learnings from the OpenROAD ProjectabstractWe describe the planned Alpha release of OpenROAD, an open-source end-to-end silicon compiler. OpenROAD will help realize the goal of "democratization of hardware design", by reducing cost, expertise, schedule and risk barriers that confront system designers today. The development of open-source, self-driving design tools is in and of itself a "moon shot" with numerous technical and cultural challenges. The open-source flow incorporates a compatible open-source set of tools that span logic synthesis, floorplanning, placement, clock tree synthesis, global routing and detailed routing. The flow also incorporates analysis and support tools for static timing analysis, parasitic extraction, power integrity analysis, and cloud deployment. We also note several observed challenges, or "lessons learned", with respect to development of open-source EDA tools and flows. Tutu Ajayi, Vidya A. Chhabria, Mateus Fogaça, Soheil Hashemi, Abdelrahman Hosny, Andrew B. Kahng, Jeongsup Lee, Uday Mallappa, Marina Neseem, Geraldo Pradipta, Sherief Reda, Mehdi Saligane, Sachin S. Sapatnekar, Carl Sechen, Mohamed Shalan, William Swartz, Lutong Wang, Zhehong Wang, Mingyu Woo, Bangqi Xu |
DAC | 6 |
| 2019 | Detailed Placement for IR Drop Mitigation by Power Staple Insertion in Sub-10nm VLSIabstractPower Delivery Network (PDN) is one of the most challenging topics in modern VLSI design. Due to aggressive technology node scaling, resistance of back-end-of-line (BEOL) layers increases dramatically in sub-10nm VLSI, causing high supply voltage (IR) drop. To solve this problem, pre-placed or post-placed power staples are inserted in pin-access layers to connect adjacent power rails and reduce PDN resistance, at the cost of reduced routing flexibility, or reduced power staple insertion opportunity. In this work, we propose dynamic programming-based single-row and double-row detailed placement optimizations to maximize the power staple insertion in a post-placement flow. We further propose metaheuristics to improve the quality of result. Compared to the traditional post-placement flow, we achieve up to 13.2% (10mV ) reduction in IR drop, with almost no WNS degradation. Sun ik Heo, Andrew B. Kahng, Lutong Wang, Chutong Yang |
DATE | 2 |
| 2019 | Power Delivery Pathfinding for Emerging Die-to-Wafer Integration TechnologyabstractIn advanced technology nodes, emerging die-to-wafer (D2W) integration technology is a promising "More Than Moore" lever for continued scaling of system capability and value. In D2W 3D IC implementation, the power delivery network (PDN) is crucial to meeting design specifications. However, determining the optimal PDN design is nontrivial. On the one hand, to meet the IR drop requirement, denser power mesh is desired. On the other hand, to meet the timing requirement for a high-utilization design, more routing resource should be available for signal routing. Moreover, additional competition between signal routing and power routing is caused by inter-tier vertical interconnects in 3D IC. In this paper, we propose a power delivery pathfinding methodology for emerging die-to-wafer integration, which seeks to identify an optimal or near-optimal PDN for a given design and PDN specification. Our pathfinding methodology exploits models for routability and worst IR drop, which helps reduce iterations between PDN design and circuit design in 3D IC implementation. We present validations with real design examples and a 28nm foundry technology. Andrew B. Kahng, Seokhyeong Kang, Seungwon Kim, Kambiz Samadi, Bangqi Xu |
DATE | 1 |
| 2019 | "Unobserved Corner" Prediction: Reducing Timing Analysis Effort for Faster Design Convergence in Advanced-Node DesignabstractWith diminishing margins for leading-edge products in advanced technology nodes, design closure and accuracy of timing analysis have emerged as serious concerns. A significant portion of design turnaround time is spent on timing analysis at combinations of process, voltage and temperature (PVT) corners. At the same time, accurate, signoff-quality timing analysis is desired during place-and-route and optimization steps, to avoid loops in the flow as well as overdesign that wastes area and power. We observe that timing results for a given path at different corners will have strong correlations, if only as a consequence of physics of devices and interconnects. We investigate a data-driven approach, based on multivariate linear regression, to predict the timing analysis at unobserved corners from analysis results at observed corners. We use a simple backward stepwise selection strategy to choose which corners to observe and which to predict. In order to accelerate convergence of the design process, the model must yield predicted values (from analysis at a limited number of observed corners) that are sufficiently accurate to substitute for unobserved ones. Our empirical results indicate that this is likely the case. With a 1M-instance example in foundry 16nm enablement, we obtain a model based on 10 observed corners that predicts timing results at the remaining 48 unobserved corners with less than 0.5% relative root mean squared error, and 99% of the model's relative prediction errors are less than 0.6%. Andrew B. Kahng, Uday Mallappa, Lawrence K. Saul, Shangyuan Tong |
DATE | 1 |
| 2019 | DATC RDF-2019: Towards a Complete Academic Reference Design FlowabstractWe describe a new RDF-2019 release of the IEEE CEDA DATC Robust Design Flow (RDF). RDF-2019 enhances the DATC RDF to span the entire RTL-to-GDS IC implementation flow, from logic synthesis to detailed routing. The new release represents a significant revision of the previously-reported RDF-2018 flow. Noteworthy vertical extensions include addition of logic synthesis starting from pure behavioral RTL Verilog RTL; floorplanning that includes initial DEF creation, I/O placement and PDN layout generation; and clock tree synthesis between placement legalization and global routing. A number of horizontal extensions to RDF are achieved by incorporating additional tool options at the static timing analysis, global placement, gate sizing, and detailed routing stages of the flow. Further, for the first time, multiple open-source realizations of the entire RDF tool chain are available. Last, RDF-2019 provides significantly enhanced support of and interoperability with industry-standard tools and design formats (LEF/DEF, SPEF, Liberty, SDC, etc.). We illustrate the configuration and use of RDF-2019, with example results on open as well as commercial design enablements. Jianli Chen, Iris Hui-Ru Jiang, Jinwook Jung, Andrew B. Kahng, Victor N. Kravets, Yih-Lang Li, Shih-Ting Lin, Mingyu Woo |
ICCAD | 4 |
| 2019 | IncPIRD: Fast Learning-Based Prediction of Incremental IR DropabstractThe on-chip power delivery network (PDN) is an essential element of physical implementation that strongly determines functionality, quality and reliability of a given IC product. To meet IR drop requirements, a denser power grid is desirable. On the other hand, to meet timing and layout density requirements, a sparser power grid leaves more resources for routing. Often, numerous time-consuming iterations among PDN design, IR analysis, and floorplanning or placement are needed during the physical implementation of modern high-performance designs. Thus, fast and accurate incremental IR prediction has emerged as a critical need, as it can potentially reduce the turnaround time between design and analysis and help improve design convergence. In this work, we apply superposition and partitioning techniques to extract relevant electrical features of a given SOC floorplan and PDN. We then use a machine learning model to predict the updated static IR drop for each power node (having tap current source attached) in the design throughout a series of changes (PDN modification, block movement, block power change, power pad movement) to the SOC floorplan, without needing to rerun a golden IR drop tool. We develop our model with more than 150 generated SOC floorplans with different PDN structures in 28nm foundry technology. Compared to an industry-leading, golden IR drop signoff tool (ANSYS RedHawk), we achieve 20-1000× speedup with less than 1 mV average absolute error and approximately 5m V maximum absolute error. Chia-Tung Ho, Andrew B. Kahng |
ICCAD | 2 |
| 2019 | Looking Into the Mirror of Open Source: Invited PaperabstractThe DARPA IDEA program has brought unprecedented resources and attention to development of open-source EDA. As this session convenes, IDEA is well into its second year, with “alpha” releases in the rear-view mirror, and a public v1.0 unveiling just eight months away. This talk will give a perspective on recent open-source development for digital layout generation. First, any technology should have a roadmap of quantifiable metrics, requirements, and potential solutions - projected on the axis of time. What are key aspects of such a roadmap for open-source EDA technology? Second, what does open-source EDA tell us about ourselves? The goal of open-source is a mirror for the ecosystem of academic research, semiconductor design, and commercial EDA. Who is willing to contribute? Who is capable of contributing? What are bars for “moving the needle” and sustainability? Andrew B. Kahng |
ICCAD | 1 |
| 2019 | Enhancing sensitivity-based power reduction for an industry IC design context
Hamed Fatemi, Andrew B. Kahng, Hyein Lee 0001, Jiajia Li 0002, José Pineda de Gyvez |
Integr. | 2 |
| 2019 | RePlAce: Advancing Solution Quality and Routability Validation in Global PlacementabstractThe Nesterov's method approach to analytic placement has recently demonstrated strong solution quality and scalability. We dissect the previous implementation strategy and show that solution quality can be significantly improved using two levers: 1) constraint-oriented local smoothing and 2) dynamic step size adaptation. We propose a new density function that comprehends local overflow of area resources; this enables a constraint-oriented local smoothing at per-bin granularity. Our improved dynamic step size adaptation automatically determines step size and effectively allocates optimization effort to significantly improve solution quality without undue runtime impact. Our resulting global placement tool, RePlAce, achieves an average of 2.00% half-perimeter wirelength (HPWL) reduction over all best known ISPD-2005 and ISPD-2006 benchmark results, and an average of 2.73% over all best known modern mixed-size (MMS) benchmark results, without any benchmark-specific code or tuning. We further extend our global placer to address routability, and achieve on average 8.50%-9.59% scaled HPWL reduction over previous leading academic placers for the DAC-2012 and ICCAD-2012 benchmark suites. To our knowledge, RePlAce is the first work to achieve superior solution quality across all the ISPD-2005, ISPD-2006, MMS, DAC-2012, and ICCAD-2012 benchmark suites with a single global placement engine. Chung-Kuan Cheng, Andrew B. Kahng, Ilgweon Kang, Lutong Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Enhanced Optimal Multi-Row Detailed Placement for Neighbor Diffusion Effect Mitigation in Sub-10 nm VLSIabstractLayout-dependent effect causes variation in device performance as well as mismatch in model-hardware correlation in sub-10 nm nodes. In order to effectively explore the power-performance envelope for IC design, cell libraries must provide cells with different diffusion heights, leading to neighbor diffusion effect (NDE) due to inter-cell diffusion height change (diffusion steps). Special filler cells can protect against steps to functional cells, but with nontrivial area overhead. In this paper, we develop dynamic programming (DP)-based single-row and multi-row (MR) detailed placement optimizations that optimally reduce inter-cell diffusion steps to mitigate the impacts of NDE. Compared to previous works, our algorithms are capable of exploring richer solution spaces as they support cell flipping, relocating, and reordering across cell rows; we also consider cell displacement, flipping, and wirelength costs. Notably, to our knowledge, our MR DP-based optimization algorithm is the first to optimally handle inter-row cell relocating and reordering. We also explore various metaheuristic configurations to further improve the solution quality. Last, we develop a timing-aware approach, which is capable of creating intentional steps that can potentially improve the drive strength of critical cells. The optimality is in terms of maximum diffusion step reduction, for given displacement range, reordering range and cell variants. Additionally, the above range definitions, including the definition of the ordering of cells, depend on assumptions described in Sections IV and V. Changho Han, Andrew B. Kahng, Lutong Wang, Bangqi Xu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | New directions for learning-based IC design tools and methodologiesabstractDesign-based equivalent scaling now bears much of the burden of continuing the semiconductor industry's trajectory of Moore's-Law value scaling. In the future, reductions of design effort and design schedule must comprise a substantial portion of this equivalent scaling. In this context, machine learning and deep learning in EDA tools and design flows offer enormous potential for value creation. Examples of opportunities include: improved design convergence through prediction of downstream flow outcomes; margin reduction through new analysis correlation mechanisms; and use of open platforms to develop learning-based applications. These will be the foundations of future design-based equivalent scaling in the IC industry. This paper describes several near-term challenges and opportunities, along with concrete existence proofs, for application of learning-based methods within the ecosystem of commercial EDA, IC design, and academic research. Andrew B. Kahng |
ASP-DAC | 1 |
| 2018 | Reducing time and effort in IC implementation: a roadmap of challenges and solutionsabstractTo reduce time and effort in IC implementation, fundamental challenges must be solved. First, the need for (expensive) humans must be removed wherever possible. Humans are skilled at predicting downstream flow failures, evaluating key early decisions such as RTL floorplanning, and deciding tool/flow options to apply to a given design. Achieving human-quality prediction, evaluation and decision-making will require new machine learning-centric models of both tools and designs. Second, to reduce design schedule, focus must return to the long-held dream of single-pass design. Future design tools and flows that never require iteration (i.e., that never fail, but without undue conservatism) demand new paradigms and core algorithms for parallel, cloud-based design automation. Third, learning-based models of tools and flows must continually improve with additional design experiences. Therefore, the EDA and design ecosystem must develop new infrastructure for ML model development and sharing. Andrew B. Kahng |
DAC | 1 |
| 2018 | Leveraging thermally-aware chiplet organization in 2.5D systems to reclaim dark siliconabstractAs on-chip power densities of manycore systems continue to increase, one cannot simultaneously run all the cores due to thermal constraints. This phenomenon, known as the `dark silicon' problem, leads to inactive regions on the chip and limits the performance of manycore systems. This paper proposes to reclaim dark silicon through a thermally-aware chiplet organization technique in 2.5D manycore systems. The proposed technique adjusts the interposer size and the spacing between adjacent chiplets to reduce the peak temperature of the overall system. In this way, a system can operate with a larger number of active cores at a higher frequency without violating thermal constraints, thereby achieving higher performance. To determine the chiplet organization that jointly maximizes performance and minimizes manufacturing cost, we formulate and solve an optimization problem that considers temperature and interposer size constraints of 2.5D systems. We design a multi-start greedy approach to find (near-)optimal solutions efficiently. Our analysis demonstrates that by using our proposed technique, an optimized 2.5D manycore system improves performance by 41% and 16% on average and by up to 87% and 39% for temperature thresholds of 85°C and 105°C, respectively, compared to a traditional single-chip system at the same manufacturing cost. When maintaining the same performance as an equivalent single-chip system, our approach is able to reduce the 2.5D system manufacturing cost by 36%. Furkan Eris, Ajay Joshi, Andrew B. Kahng, Yenai Ma, Saiful A. Mojumder, Tiansheng Zhang |
DATE | 3 |
| 2018 | A cross-layer methodology for design and optimization of networks in 2.5D systemsabstract2.5D integration technology is gaining popularity in the design of homogeneous and heterogeneous many-core computing systems. 2.5D network design, both inter- and intra-chiplet, impacts overall system performance as well as its manufacturing cost and thermal feasibility. This paper introduces a cross-layer methodology for designing networks in 2.5D systems. We optimize the network design and chiplet placement jointly across logical, physical, and circuit layers to achieve an energy-efficient network, while maximizing system performance, minimizing manufacturing cost, and adhering to thermal constraints. In the logical layer, our co-optimization considers eight different network topologies. In the physical layer, we consider routing, microbump assignment, and microbump pitch constraints to account for the extra costs associated with microbump utilization in the inter-chiplet communication. In the circuit layer, we consider both passive and active links with five different link types, including a gas station link design. Using our cross-layer methodology results in more accurate determination of (superior) inter-chiplet network and 2.5D system designs compared to prior methods. Compared to 2D systems, our approach achieves 29% better performance with the same manufacturing cost, or 25% lower cost with the same performance. Ayse K. Coskun, Furkan Eris, Ajay Joshi, Andrew B. Kahng, Yenai Ma, Vaishnav Srinivas |
ICCAD | 4 |
| 2018 | TritonRoute: an initial detailed router for advanced VLSI technologiesabstractDetailed routing is a dead-or-alive critical element in design automation tooling for advanced node enablement. However, very few works address detailed routing in the recent open literature, particularly in the context of modern industrial designs and a complete, end-to-end flow. The ISPD-2018 Initial Detailed Routing Contest addressed this gap for modern industrial designs, using a reduced design rules set. In this work, we present TritonRoute, an initial detailed router for the ISPD-2018 contest. Given route guides from global routing, the initial detailed routing stage should generate a detailed routing solution honoring the route guides as much as possible, while minimizing wirelength, via count and various design rule violations. In our work, the key contribution is intra-layer parallel routing, where we partition each layer into parallel panels and route each panel using an Integer Linear Programming-based algorithm. We sequentially route layer by layer from the bottom to the top. We evaluate our router using the official ISPD-2018 benchmark suite and show that we reduce the contest metric by up to 74%, and on average 50%, compared to the first-place routing solution for each testcase. Andrew B. Kahng, Lutong Wang, Bangqi Xu |
ICCAD | 1 |
| 2018 | Using Machine Learning to Predict Path-Based Slack from Graph-Based Timing AnalysisabstractWith diminishing margins in advanced technology nodes, accuracy of timing analysis is a serious concern. Improved accuracy helps to reduce overdesign, particularly in P&R-based optimization and timing closure steps, but comes at the cost of runtime. A major factor in accurate estimation of timing slack, especially for low-voltage corners, is the propagation of transition time. In graph-based analysis (GBA), worst-case transition time is propagated through a given gate, independent of the path under analysis, and is hence pessimistic. The timing pessimism results in overdesign and/or inability to fully recover power and area during optimization. In path-based analysis (PBA), pathspeci?c transition times are propagated, reducing pessimism. However, PBA incurs severe (4X or more) runtime overheads relative to GBA, and is often avoided in the early stages of physical implementation. With numerous operating corners, use of PBA is even more burdensome. In this paper, we propose a machine learning model, based on bigrams of path stages, to predict expensive PBA results from relatively inexpensive GBA results. We identify electrical and structural features of the circuit that affect PBA-GBA divergence with respect to endpoint arrival times. We use GBA and PBA analysis of a given testcase design along with arti?cially generated timing paths, in combination with a classi?cation and regression tree (CART) approach, to develop a predictive model for PBA-GBA divergence. Empirical studies demonstrate that our model has the potential to substantially reduce pessimism while retaining the lower turnaround time of GBA analysis. For example, a model trained on a post-CTS and tested on a post-route database for the leon3mp design in 28nm FDSOI foundry enablement reduces divergence from true PBA slack (i.e., model prediction divergence, versus GBA divergence) from 9.43ps to 6.06ps (mean absolute error), 26.79ps to 19.72ps (99th percentile error), and 50.78ps to 39.46ps (maximum error). Andrew B. Kahng, Uday Mallappa, Lawrence K. Saul |
ICCD | 1 |
| 2018 | Prim-Dijkstra Revisited: Achieving Superior Timing-driven Routing TreesabstractThe Prim-Dijkstra (PD ) construction [1] was first presented over 20 years ago as a way to efficiently trade off between shortest-path and minimum-wirelength routing trees. This approach has stood the test of time, having been integrated into leading semiconductor design methodologies and electronic design automation tools. PD optimizes the conflicting objectives of wirelength (WL) and source-sink pathlength (PL) by blending the classic Prim and Dijkstra spanning tree algorithms. However, as this work shows, PD can sometimes demonstrate significant suboptimality for both WL and PL. This quality degradation can be especially costly for advanced nodes because (i) wire delays form a much larger component of total stage delay, i.e., timing-driven routing is critical, and (ii) modern designs are severely power-constrained (e.g., mobile, IoT), which makes low-capacitance wiring important. Consequently, achieving a good timing and power tradeoff for routing is required to build a market-leading product[2]. This work introduces a new problem formulation that incorporates the total detour cost in the objective function to optimize the detour to every sink in the tree, not just the worst detour. We then propose a new PD-II construction which directly improves upon the original PD construction by repairing the tree to simultaneously reduce both WL and PL. The PD-II approach achieves improvement for both objectives, making it a clear win over PD, for virtually zero additional runtime cost. PD-II is a spanning tree algorithm (which is useful for seeding global routing); however, since Steiner trees are needed for timing estimation, this work also includes a post-processing algorithm called DAS to convert PD-II trees into balanced Steiner trees. Experimental results demonstrate that this construction outperforms the recent state-of-the-art academic tool, SALT [36], for high-fanout nets, achieving up to 36.46% PL improvement with similar WL on average for 20K nets of size ≥ 32 terminals from DAC 2012 contest benchmark designs [37]. Charles J. Alpert, Wing-Kai Chow, Kwangsoo Han, Andrew B. Kahng, Zhuo Li 0001, Derong Liu 0002, Sriram Venkatesh |
ISPD | 4 |
| 2018 | Theory and Algorithms of Physical DesignabstractNo abstract available. Chung-Kuan Cheng, T. C. Hu, Andrew B. Kahng |
ISPD | 3 |
| 2018 | Machine Learning Applications in Physical Design: Recent Results and DirectionsabstractIn the late-CMOS era, semiconductor and electronics companies face severe product schedule and other competitive pressures. In this context, electronic design automation (EDA) must deliver "design-based equivalent scaling" to help continue essential industry trajectories. A powerful lever for this will be the use of machine learning techniques, both inside and "around" design tools and flows. This paper reviews opportunities for machine learning with a focus on IC physical implementation. Example applications include (1) removing unnecessary design and modeling margins through correlation mechanisms, (2) achieving faster design convergence through predictors of downstream flow outcomes that comprehend both tools and design instances, and (3) corollaries such as optimizing the usage of design resources licenses and available schedule. The paper concludes with open challenges for machine learning in IC physical design. Andrew B. Kahng |
ISPD | 1 |
| 2018 | Influence of Professor T. C. Hu's Works on Fundamental Approaches in LayoutabstractProfessor T. C. Hu has made numerous pioneering and fundamental contributions in combinatorial algorithms, mathematical programming and operations research. His seminal 1985 IEEE book VLSI Circuit Layout: Theory and Design, coedited with Prof. E. S. Kuh, shaped algorithmic and optimization perspectives, as well as basic frameworks, for IC physical design throughout the following decades [44]. Indeed, Professor Hu gave a keynote address at the very first International Symposium on Physical Design, in April 1997 [41]. His research approach of (i) studying small and/or extremal cases first, (ii) always seeking to establish error bounds or other properties of heuristic outcomes (hence, always seeking to understand optimal solutions), (iii) approaching new problems from as fresh a perspective as possible, and (iv) pursuing simplicity and beauty in both formulation and exposition, has influenced generations of researchers. This paper complements [20] in recounting highlights of Professor Hu's contributions to, and impacts on, physical design. The focus is on several problems of "connection", as well as the concept of "shadow price", as they relate to layout. Andrew B. Kahng |
ISPD | 1 |
| 2018 | Wot the L: Analysis of Real versus Random Placed Nets, and Implications for Steiner Tree HeuristicsabstractThe NP-hard Rectilinear Steiner Minimum Tree (RSMT) problem has been studied in the VLSI physical design literature for well over three decades. Fast estimators of RSMT cost (which reflects routed wirelength) are a required ingredient of modern physical planning and global placement methods. Constructive estimators build heuristic RSMTs whose costs are used as wirelength estimates; notably, these include FLUTE[8]. Analytic and lookup table-based estimators include the methods of Cheng[7] and Caldwell et al. [3]; the latter is based on both the number of points and the aspect ratio of the pointset in the RSMT instance. We observe that the physical design literature has numerous evaluations of RSMT heuristics and estimators on random pointsets, and that the relative merits of heuristics and estimators have been determined based on this use of random pointsets. In this paper, we show that a pointset attribute which we call L-ness highlights the difference between real placements and random placements of net pins. We explain why placements of netlists in practice result in pointsets with much higher L-ness than random pointsets, and we confirm this difference empirically for both academic and commercial placement tools. We further present an improved lookup table-based RSMT cost estimator that includes an L-ness parameter. Last, we illustrate how differences between Steiner tree heuristics can change depending on whether real or random pointsets are used in the evaluation. Andrew B. Kahng, Christopher Moyes, Sriram Venkatesh, Lutong Wang |
ISPD | 1 |
| 2018 | Silicon Photonics for Computing SystemsabstractNo abstract available. Jiang Xu 0001, Yuichi Nakamura 0002, Andrew B. Kahng |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2018 | Design Implementation With Noninteger Multiple-Height Cells for Improved Design Quality in Advanced NodesabstractStandard-cell libraries can be developed with different cell heights (e.g., in FinFET technology, corresponding to different numbers of fins). Larger cell heights provide higher drive strengths, but at the cost of larger area and power consumption as well as pin capacitance. Cells with smaller heights are relatively smaller in area, but have weaker drive strengths and are more likely to suffer from routing congestion and pin accessibility issues. Existing design methodologies and tool flows are able to mix cells with different heights at the block level (i.e., each block contains cells with heights being integer multiples of a particular “single row” cell height). To our knowledge, no design methodology in the literature or in production mixes cells with different, noninteger multiple heights in a fine-grained manner. In this paper, we propose a novel physical design optimization flow to implement design blocks with mixed cell heights in a fine-grained manner. Our optimization resolves the “chicken-and-egg” loop between floorplan site definition and the optimized choices of cell heights after placement with full comprehension of the constraints and costs of mixing cells of different heights (e.g., the “breaker cell” area overheads of row alignment between sub-blocks of 8 T and 12 T cell rows), our optimization achieves up to over 30% area and power reductions versus 12 T-only implementation while maintaining the same performance, and up to over 10% performance improvement along with power and area reductions versus 8 T-only implementation. Sorin Dobre, Andrew B. Kahng, Jiajia Li 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | PROBE: A Placement, Routing, Back-End-of-Line Measurement UtilityabstractIn advanced technology nodes, correctly choosing among available back-end-of-line (BEOL) stack options is important to meet stringent design quality of results requirements. However, it is nontrivial to evaluate BEOL stack options since the routing outcomes highly depend on the input design (e.g., netlist, placement, etc.). In this paper, we propose a systematic framework to measure routing capacity of a BEOL stack as well as inherent capability of routers. Based on our experimental results, we observe consistent results across mesh-like placement and placements from various placers. Also, our proposed framework enables new insights into important questions regarding BEOL stack options. Using our framework, we further study the relation between the routing hotspot size and routing failure empirically. Lastly, we present an analytical study based on exponentiation of a Markov transition matrix about the impact of design size on routing failure. Alex Kahng, Andrew B. Kahng, Hyein Lee 0001, Jiajia Li 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | Floorplan and placement methodology for improved energy reduction in stacked power-domain designabstractEnergy and battery lifetime constraints are critical challenges to IC designs. Stacked power-domain implementation, which stacks voltage domains in a design, can effectively improve the power delivery efficiency and thus improve battery lifetime. However, such an approach requires balanced current between different domains across multiple operating scenarios. Furthermore, level shifter insertion (together with shifters' delay impacts), along with placement constraints imposed by power domain regions, can incur power and area penalties. To our knowledge, no existing work performs sub-block-level partitioning optimization for stacked-domain designs. In this paper, we present an optimization framework for stacked-domain designs. Based on an initial placement solution, we apply a flow-based partitioning that is aware of multiple operating scenarios, cell placement, and timing-critical paths to partition cells into two power domains with balanced current and minimized number of inserted level shifters. We further propose heuristics to define regions for each power domain so as to minimize placement perturbation, as well as a dynamic programming-based method to minimize the area cost of power domain generation. In an updated floorplan, we perform matching-based optimization to insert level shifters with minimized wirelength penalty. Overall, our method achieves more than ~10% and 3X battery lifetime improvements in function and sleep modes, respectively. Kristof Blutman, Hamed Fatemi, Andrew B. Kahng, Ajay Kapoor, Jiajia Li 0002, José Pineda de Gyvez |
ASP-DAC | 3 |
| 2017 | Vertical M1 Routing-Aware Detailed Placement for Congestion and Wirelength Reduction in Sub-10nm NodesabstractAggressive pitch scaling in sub-10nm nodes has introduced complex design rules which make routing extremely challenging. Cell architectures have also been changed to meet the design rules. For example, metal layers below M1 are used to gain additional routing resources. New cell architectures wherein inter-row M1 routing is allowed force consideration of vertical alignment of cells. In this work, we propose a mixed-integer linear programming (MILP)-based, detailed placement optimization to maximize direct vertical M1 routing utilization for congestion and wirelength reduction. Peter Debacker, Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001, Praveen Raghavan, Lutong Wang |
DAC | 3 |
| 2017 | Optimal multi-row detailed placement for yield and model-hardware correlation improvements in sub-10nm VLSIabstractIn sub-10nm, nodes, a change or step in diffusion height between adjacent standard cells causes yield loss as well as a form of model-hardware miscorrelation called neighbor diffusion effect (NDE). Cell libraries must inevitably have multiple diffusion heights (numbers of fins in PFETs and NFETs) in order to enable flexible exploration of the power-performance envelope for design. However, this brings step-induced risks of NDE, for which guardbanding is costly, as well as yield loss. Special filler cells can protect against harmful NDE effects, but are costly in terms of area. In this work, we develop dynamic programming-based single-row and double-row detailed placement optimizations that optimally minimize the impacts of NDE. Our algorithms support a richer set of cell movements than in previous works - i.e., flipping, relocating and reordering within the original row; we also consider cell displacement and flipping costs. Importantly, to our knowledge, our dynamic programming-based optimal detailed placement algorithm is the first to handle multiple rows with multiple-height cells that can be reordered. We further develop a timing-aware approach, which is capable of recovering (or, improving) the worst negative slack (WNS) by creating additional diffusion steps around timing-critical cells. Changho Han, Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001, Lutong Wang, Bangqi Xu |
ICCAD | 3 |
| 2017 | ILP-Based Identification of Redundant Logic Insertions for Opportunistic Yield Improvement during Early Process LearningabstractAs semiconductor technology advances, leading-edge product companies must rapidly improve yield for designs that seek to reach mass production while early in the adoption of a new technology node; otherwise, products may be unviable in the marketplace. In this paper, we are the first to study the possible mitigation by opportunistic, last-stage redundant logic insertion to mitigate yield loss in early advanced-node production. We describe a yield estimation methodology, and propose an integer linear programming (ILP)-based optimization of redundant logic insertion for opportunistic yield optimization. In our approach, we first identify potential logic clusters for replication by top-down application of multilevel FM partitioning. We then select promising clusters whose replication maximizes design yield without hurting design timing. Our experimental results show the potential for significant yield improvement with minor timing impact. Tuck-Boon Chan, Wei-Ting Jonas Chan, Andrew B. Kahng |
ICCD | 3 |
| 2017 | Routability Optimization for Industrial Designs at Sub-14nm Process Nodes Using Machine LearningabstractDesign rule check (DRC) violations after detailed routing prevent a design from being taped out. To solve this problem, state-of-the-art commercial EDA tools global-route the design to produce a global-route congestion map; this map is used by the placer to optimize the placement of the design to reduce detailed-route DRC violations. However, in sub-14nm processes and beyond, DRCs arising from multiple patterning and pin-access constraints drastically weaken the correlation between global-route congestion and detailed-route DRC violations. Hence, the placer|based on the global-route congestion map|may leave too many detailed-route DRC violations to be fixed manually by designers. In this paper, we present a method that employs (1) machine-learning techniques to effectively predict detailed-route DRC violations after global routing and (2) detailed placement techniques to effectively reduce detailed-route DRC violations. We demonstrate on several layouts of a sub-14nm industrial design that this method predicts the locations of 74% of the detailed-route DRCs (with false positive prediction rate below 0.2%) and automatically reduces the number of detailed-route DRC violations by up to 5x. Whereas previous works on machine learning for routability [30] [4] have focused on routability prediction at the floorplanning and placement stages, ours is the first paper that not only predicts the actual locations of detailed-route DRC violations but furthermore optimizes the design to significantly reduce such violations. Wei-Ting Jonas Chan, Pei-Hsin Ho, Andrew B. Kahng, Prashant Saxena |
ISPD | 3 |
| 2017 | Revisiting 3DIC benefit with multiple tiers
Wei-Ting Jonas Chan, Andrew B. Kahng, Jiajia Li 0002 |
Integr. | 2 |
| 2017 | Trading Accuracy for Energy in Stochastic Circuit DesignabstractAs we approach the limits of traditional Moore’s-Law scaling, alternative computing techniques that consume energy more efficiently become attractive. Stochastic computing (SC), as a re-emerging computing technique, is a low-cost and error-tolerant alternative to conventional binary circuits in several important applications such as image processing and communications. SC allows a natural accuracy-energy tradeoff that has been exploited in the past. This article presents an accuracy-energy tradeoff technique for SC circuits that reduces their energy consumption with virtually no accuracy loss. To this end, we employ voltage or frequency scaling, which normally reduce energy consumption at the cost of timing errors. Then we show that due to their inherent error tolerance, SC circuits operate satisfactorily without significant accuracy loss even with aggressive scaling. This significantly improves their energy efficiency. In contrast, conventional binary circuits quickly fail as the supply voltage decreases. To find the most energy-efficient operating point of an SC circuit, we propose an error estimation method that allows us to quickly explore the circuit’s design space. The error estimation method is based on Markov chain and least-squares regression. Furthermore, we investigate opportunities to optimize SC circuits under such aggressive scaling. We find that logical and physical design techniques can be combined to significantly expand the already-powerful accuracy-energy tradeoff possibilities of SC. In particular, we demonstrate that careful adjustment of path delays can lead to significant error reduction under voltage and frequency scaling. We perform buffer insertion and route detouring to achieve more balanced path delays. These techniques differ from conventional path-balancing techniques whose goal is to minimize power consumption by resizing the non-critical paths. The goal of our path-balancing approach is to increase error cancellation chances in voltage-/frequency-scaled SC circuits. Our circuit optimization comprehends the tradeoff between power overheads due to inserted buffers and wires versus the energy reduction from supply voltage downscaling enabled by more balanced path delays. Simulation results show that our optimized SC circuits can tolerate aggressive voltage scaling with no significant signal-to-noise ratio (SNR) degradation. In one example, a 40% supply voltage reduction (1V to 0.6V) on the SC circuit leads to 66% energy saving (20.7pJ to 6.9pJ) and makes it more efficient than its conventional binary counterpart. In the same example, a 100% frequency boosting (400ps to 200ps) of the optimized circuits leads to no significant SNR degradation. We also show that process variation and temperature variation have limited impact on optimized SC circuits. The error change is less than 5% when temperature changes by 100°C or process condition changes from worst case to best case. Armin Alaghi, Wei-Ting Jonas Chan, John P. Hayes, Andrew B. Kahng, Jiajia Li 0002 |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2017 | CACTI 7: New Tools for Interconnect Exploration in Innovative Off-Chip MemoriesabstractHistorically, server designers have opted for simple memory systems by picking one of a few commoditized DDR memory products. We are already witnessing a major upheaval in the off-chip memory hierarchy, with the introduction of many new memory products—buffer-on-board, LRDIMM, HMC, HBM, and NVMs, to name a few. Given the plethora of choices, it is expected that different vendors will adopt different strategies for their high-capacity memory systems, often deviating from DDR standards and/or integrating new functionality within memory systems. These strategies will likely differ in their choice of interconnect and topology, with a significant fraction of memory energy being dissipated in I/O and data movement. To make the case for memory interconnect specialization, this paper makes three contributions. First, we design a tool that carefully models I/O power in the memory system, explores the design space, and gives the user the ability to define new types of memory interconnects/topologies. The tool is validated against SPICE models, and is integrated into version 7 of the popular CACTI package. Our analysis with the tool shows that several design parameters have a significant impact on I/O power. We then use the tool to help craft novel specialized memory system channels. We introduce a new relay-on-board chip that partitions a DDR channel into multiple cascaded channels. We show that this simple change to the channel topology can improve performance by 22% for DDR DRAM and lower cost by up to 65% for DDR DRAM. This new architecture does not require any changes to DIMMs, and it efficiently supports hybrid DRAM/NVM systems. Finally, as an example of a more disruptive architecture, we design a custom DIMM and parallel bus that moves away from the DDR3/DDR4 standards. To reduce energy and improve performance, the baseline data channel is split into three narrow parallel channels and the on-DIMM interconnects are operated at a lower frequency. In addition, this allows us to design a two-tier error protection strategy that reduces data transfers on the interconnect. This architecture yields a performance improvement of 18% and a memory power reduction of 23%. The cascaded channel and narrow channel architectures serve as case studies for the new tool and show the potential for benefit from re-organizing basic memory interconnects. Rajeev Balasubramonian, Andrew B. Kahng, Naveen Muralimanohar, Ali Shafiee, Vaishnav Srinivas |
ACM Trans. Archit. Code Optim. | 2 |
| 2017 | Adaptive Tuning of Photonic Devices in a Photonic NoC Through Dynamic Workload AllocationabstractPhotonic network-on-chip (PNoC) is a promising candidate to replace traditional electrical NoC in manycore systems that require substantial bandwidths. The photonic links in the PNoC comprise laser sources, optical ring resonators, passive waveguides, and photodetectors. Reliable link operation requires laser sources and ring resonators to have matching optical frequencies. However, inherent thermal sensitivity of photonic devices and manufacturing process variations can lead to a frequency mismatch. To avoid this mismatch, micro-heaters are used for thermal trimming and tuning, which can dissipate a significant amount of power. This paper proposes a novel FreqAlign workload allocation policy, accompanying an adaptive frequency tuning (AFT) policy, that is capable of reducing thermal tuning power of PNoC. FreqAlign uses thread allocation and thread migration to control temperature for matching the optical frequencies of ring resonators in each photonic link. The AFT policy reduces the remaining optical frequency difference among ring resonators and corresponding on-chip laser sources by hardware tuning methods. We use a full modeling stack of a PNoC that includes a performance simulator, a power simulator, and a thermal simulator with a temperature-dependent laser source power model to design and evaluate our proposed policies. Our experimental results demonstrate that FreqAlign reduces the resonant frequency gradient between ring resonators by 50%-60% when compared to existing workload allocation policies. Coupled with AFT, FreqAlign reduces localized thermal tuning power by 19.28 W on average, and is capable of saving up to 34.57 W when running realistic loads in a 256-core system without any performance degradation. José L. Abellán, Ayse K. Coskun, Anjun Gu, Warren Jin 0002, Ajay Joshi, Andrew B. Kahng, Jonathan Klamkin, Cristian Morales, John Recchio, Vaishnav Srinivas, Tiansheng Zhang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2017 | Benchmarking of Mask Fracturing HeuristicsabstractAggressive resolution enhancement techniques such as inverse lithography (ILT) often lead to complex, nonrectilinear mask shapes which make mask writing extremely slow and expensive. To reduce shot count of complex mask shapes, mask writers allow overlapping shots, due to which the problem of fracturing mask shapes with minimum shot count is NP-hard. The need to account for e-beam proximity effect makes mask fracturing even more challenging. Although a number of fracturing heuristics have been proposed, there has been no systematic study to analyze the quality of their solutions. In this paper, we first propose a method to generate tight upper and lower bounds for actual ILT mask shapes by formulating mask fracturing as an integer linear program and solving it using branch and price. Since the integer program requires significant computational resources to compute reasonable bounds, we propose a new method to generate benchmarks with known optimal solutions, that can be used to evaluate the suboptimality of mask fracturing heuristics. To make the generated benchmark shapes realistic, we further propose a novel automated benchmark generation method that takes any ILT shape as input and returns a benchmark shape which looks similar to the input shape and for which the optimal fracturing solution is known. Using these methods, we compare the suboptimality of four mask fracturing heuristics. Our results show that even a state-of-the-art prototype (version of) capability within a commercial EDA tool for e-beam mask shot decomposition can be suboptimal by as much as 2.6× for real ILT shapes and by 6.0× for generated benchmarks. Tuck-Boon Chan, Puneet Gupta 0001, Kwangsoo Han, Abde Ali Kagalwalla, Andrew B. Kahng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2017 | MILP-Based Optimization of 2-D Block Masks for Timing-Aware Dummy Segment Removal in Self-Aligned Multiple Patterning LayoutsabstractSelf-aligned multiple patterning, due to its low overlay error, has emerged as the leading option for 1-D gridded back-end-of-line (BEOL) in sub-14-nm nodes. To form actual routing patterns from a uniform “sea of wires,” cut masks are needed for line-end cutting or realization of space between routing segments. The line-end cutting results in nonfunctional (i.e., dummy fill) patterns that change wire capacitance, and hence design timing and power. Therefore, to remove such dummy fill patterns, extra 2-D block masks are used. However, 2-D block masks cannot remove arbitrary dummy fill patterns, due to design rule constraints on the block mask shapes. In this paper, we address the timing-aware optimization of 2-D block mask layouts under various sets of mask rules that are derived from mask patterning technology options (e.g., 193i and 193d) for foundry 7-/5-nm (N7/N5) BEOL. Our central contribution is a mixed integer linear programming (MILP) optimization that minimizes timing impact due to dummy metal segments while satisfying block mask rules and metal density constraints. We also propose a distributed optimization flow to improve the scalability. With our optimizer, we recover up to 84% of the worst negative slack impact from dummy segments, with up to 64% dummy removal rate. We further extend our MILP to a co-optimization of cut and block masks. This paper gives new insights into fundamental limits of benefit from emerging cut and block mask technology options. Peter Debacker, Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001, Praveen Raghavan, Lutong Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Optimal Scheduling and Allocation for IC Design Management and Cost ReductionabstractA large semiconductor product company spends hundreds of millions of dollars each year on design infrastructure to meet tapeout schedules for multiple concurrent projects. Resources (servers, electronic design automation tool licenses, engineers, and so on) are limited and must be shared -- and the cost per day of schedule slip can be enormous.Co-constraintsbetween resource types (e.g., one license per every two cores (threads)) and dedicated versus shareable resource pools make scheduling and allocation hard. In this article, we formulate two mixed integer-linear programs for optimal multi-project, multi-resource allocation with task precedence and resource co-constraints. Application to a real-world three-project scheduling problem extracted from a leading-edge design center of anonymized Company X shows substantial compute and license costs savings. Compared to the product company, our solution shows that the makespan of schedule of all projects can be reduced by seven days, which not only saves ∼ 2.7% of annual labor and infrastructure costs but also enhances market competitiveness. We also demonstrate the capability of scheduling over two dozen chip development projects at the design center level, subject to resource and datacenter capacity limits as well as per-project penalty functions for schedule slips. The design center ended up purchasing 600 additional servers, whereas our solution demonstrates that the schedule can be met without having to purchase any additional servers. Application to a four-project scheduling problem extracted from a leading-edge design center in a non-US location shows availability of up to ∼ 37% headcount reduction during a half-year schedule for just one type of chip design activity. Prabhav Agrawal, Mike Broxterman, Biswadeep Chatterjee, Patrick Cuevas, Kathy H. Hayashi, Andrew B. Kahng, Pranay K. Myana, Siddhartha Nath |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2017 | Logic Design Partitioning for Stacked Power DomainsabstractEnergy and battery lifetime constraints are critical challenges to IC designs. Stacked power-domain implementation, which connects voltage domains in series, can effectively improve power delivery efficiency and thus improve battery lifetime. However, such an approach requires balanced currents between different domains across multiple operating scenarios. Furthermore, level shifter insertion, along with placement constraints imposed by power domain regions, can incur significant power and area penalties. To the best of our knowledge, no existing work performs subblock-level partitioning optimization for stackeddomain designs. In this paper, we present an optimization framework for stacked-domain designs. Based on an initial placement solution, we apply a flow-based partitioning that is aware of multiple operating scenarios, cell placement, and timing-critical paths to partition cells into two power domains with balanced cross-domain current and minimized number of inserted level shifters. We further propose heuristics to define regions for each power domain so as to minimize placement perturbation, as well as a dynamic programming-based method to minimize the area cost of power domain generation. In an updated floor plan, we perform matching-based optimization to insert level shifters with minimized wirelength penalty. Overall, our method achieves an excellent current balance across stacked domains with less than 10% discrepancy, which results in up to more than 2× battery lifetime improvements. Kristof Blutman, Hamed Fatemi, Ajay Kapoor, Andrew B. Kahng, Jiajia Li 0002, José Pineda de Gyvez |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | Guest Editorial: Alternative Computing and Machine Learning for Internet of ThingsabstractThe impending Internet of Things (IoT) wave is promising to affect every aspect of our daily lives, ranging from smart things to smart buildings, smart cities, and smart environments. A lot of attention has been devoted to the tsunami of data produced by IoT, and the related means of extracting useful actionable information from it, spawning efforts in Big Data processing and machine learning. Yet, all of this does little to address the need for IoT to capture, interpret, and act on this wall of (noisy) information at the right time, at the right place, and in the right form. Conventional computing systems are a poor match to the needs of this emerging massively distributed real-time system. Hence, alternative computing techniques present an attractive alternative, trading off computational resolution for significant gains in quality-of-service energy efficiency and robustness. This observation is based on the conjecture that most applications related to IoT have an inherent error resilience and are evolutionary (that is, learning-based). Alternative computing strategies may be conceived at every level of the design hierarchy, starting from the device level with novel 3-D nonvolatile memory/logic combinations, or at the architectural level by shifting away from the traditional von Neumann architecture to different computing paradigms such as neuromorphic and/or stochastic computation all the way up to the algorithmic and data representation levels. Farshad Firouzi, Bahareh J. Farahani, Andrew B. Kahng, Jan M. Rabaey, Natasha Balac |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Delay uncertainty and signal criticality driven routing channel optimization for advanced DRAM productsabstractSignal delay uncertainty induced by crosstalk is a critical challenge to the physical design of long interconnect channels in DRAM products at the 2× and 1× technology nodes. Due to severe cost challenges in a high-volume, commodity market, layout resources including channel width, buffers, and number of metal routing layers are extremely scarce. We describe a new channel optimizer that reduces crosstalk-induced delay uncertainty, weighted by signal criticality and aware of signal activity correlations (e.g., to reduce delay uncertainty by mutual shielding). Instead of the typical signal net permutation strategy, we apply (pessimistic) timing-driven swizzling to minimize the delay uncertainty cost function. Contributions of this work include (1) an accurate and efficient analytical crosstalk delay calculator, (2) scalability up to hundreds of signals and tracks in the routing channel through use of greedy and decomposition strategies as well as a pair-swapping approach, and (3) experimental studies that demonstrate up to 24% reduction of the worst-case criticalityweighted delay uncertainty (or, 34ps of absolute delay uncertainty reduction) compared with the typical signal permutation approach. Samyoung Bang, Kwangsoo Han, Andrew B. Kahng, Mulong Luo |
ASP-DAC | 3 |
| 2016 | Learning-based prediction of embedded memory timing failures during initial floorplan designabstractEmbedded memories are critical to success or failure of complex system-on-chip (SoC) products. They can be significant yield detractors as a consequence of occupying substantial die area, creating placement and routing blockages, and having stringent Vccmin and power integrity requirements. Achieving timing-correctness for embedded memories in advanced nodes is costly (e.g., closing the design at multiple (logic-memory) cross-corners). Further, multiphysics (e.g., crosstalk, IR, etc.) signoff analyses make early understanding and prediction of timing (-correctness) even more difficult. With long tool and design closure subflow runtimes, design teams need improved prediction of embedded memory timing failures, as early as possible in the implementation flow. In this work, we propose a learning-based methodology to perform early prediction of timing failure risk given only the netlist, timing constraints, and floorplan context (wherein the memories have been placed). Our contributions include (i) identification of relevant netlist and floorplan parameters, (ii) the avoidance of long P&R tool runtimes (up to a week or even more) with early prediction, and (iii) a new implementation of Boosting with Support Vector Machine regression with focus on negative-slack outcomes through weighting in the model construction. We validate accuracy of our prediction models across a range of “multiphysics” analysis regimes, and with multiple designs and floorplans in 28FDSOI foundry technology. Our work can be used to identify which memories are “at risk”, guide floorplan changes to reduce predicted “risk”, and help refine underlying SoC implementation methodologies. Experimental results in 28nm FDSOI technology show that we can predict P&R slack with multiphysics analysis to within 253ps (average error less than 10ps) using only post-synthesis netlist, constraints and floorplan information. Our predictions are 40% more accurate than the predictions (worst-case error of 358ps and average error of 42ps) of a nonlinear Support Vector Machine model that uses only post-synthesis netlist information. Wei-Ting Jonas Chan, Kun Young Chung, Andrew B. Kahng, Nancy D. MacDonald, Siddhartha Nath |
ASP-DAC | 3 |
| 2016 | Comprehensive optimization of scan chain timing during late-stage IC implementationabstractScan chain timing is increasingly critical to test time and product cost. However, hold buffer insertions (e.g., due to large clock skew) limit scan timing improvement. Dynamic voltage drop (DVD) during scan shift further degrades scan shift timing, inducing "false failures" in silicon. Hence, new optimizations are needed in late stages of implementation when accurate (skew, DVD) information is available. We propose skew-aware scan ordering to minimize hold buffers, and DVD-aware gating insertion to improve scan shift timing slacks. Our optimizations at the post-CTS and post-routing stages reduce hold buffers by up to 82%, and DVD-induced timing degradation by up to 58%, with negligible area and power overheads. Kun Young Chung, Andrew B. Kahng, Jiajia Li 0002 |
DAC | 2 |
| 2016 | Cross-layer floorplan optimization for silicon photonic NoCs in many-core systems
Ayse K. Coskun, Anjun Gu, Warren Jin 0002, Ajay Joshi, Andrew B. Kahng, Jonathan Klamkin, Yenai Ma, John Recchio, Vaishnav Srinivas, Tiansheng Zhang |
DATE | 5 |
| 2016 | Improved performance of 3DIC implementations through inherent awareness of mix-and-match die stacking
Kwangsoo Han, Andrew B. Kahng, Jiajia Li 0002 |
DATE | 2 |
| 2016 | Measuring progress and value of IC implementation technologyabstractOver the past decade, “Moore's Law” has become increasingly well-understood as being a law of “value scaling”: success of new electronics- and semiconductor-based products depends on improved cost-efficiency, utility, and value. Design Automation (DA) provides fundamental tools and methodologies that glue together disparate technological advances - across architectures, circuits, process and integration - into actually-realized product and system benefits. However, it is still unknown how to measure and credit “progress” of DA with respect to realized product- and system-level benefits. Thus, it is also challenging for industry, academia and funding entities to envision, and define, R&D objectives and prioritizations. In this paper, we contend that “measuring progress and value” is tractable at the level of EDA technology and R&D efforts. We describe four example assessments of progress and value of IC implementation technology: (i) assessment of progress of EDA tools (e.g., for P&R and STA) across multiple releases; (ii) establishing “upper bounds” on future progress (e.g., for 3DIC layout); (iii) robust rank-ordering of alternative design enablements (e.g., including routing tools and interconnect stack options); and (iv) lowering barriers to measuring progress and value of academic research results in “more real-world” contexts. Andrew B. Kahng, Hyein Lee 0001, Jiajia Li 0002 |
ICCAD | 1 |
| 2016 | Improved flop tray-based design implementation for power reductionabstractClock network power reduction is critical in modern SoC designs. Application of flop trays (i.e., multi-bit flip-flops) can significantly reduce the number of sinks in a clock network, and thus reduce the number of clock buffers, clock wirelength, and clock network power. Shared inverters within flop trays also reduce power at the flip-flop level. However, large-size flop trays typically induce placement and routing congestion, and impose additional placement constraints on their fanin/fanout logic cones; this results in power overheads on datapaths. At the same time, to our knowledge, few previous works have studied flop trays with more than four bits. The “chicken-and-egg” loop between flop tray generation and placement optimization is another challenge to flop tray-based design [7]. In this work, we propose an optimization flow to generate and place flop trays from a library of arbitrary given sizes and aspect ratios (ARs), to achieve clock network power reduction. Our optimization starts with an initial placement solution using only single-bit flops. It then performs capacitated K-means clustering to generate solutions with different flop tray sizes and ARs. More specifically, we iteratively use (i) min-cost flow to cluster flops, and (ii) a linear programming-based optimization to determine locations of the generated flop trays. Last, we formulate an integer linear program to select the best combination of flop tray solutions (i.e., sizes and placements) with minimum displacement and number of isolated sinks. Our optimization is aware of flop tray sizes and ARs, as well as timing-critical start-end pairs. Results in foundry 28FDSOI technology show up to 32% total block power reduction as compared to designs using only single-bit flops, up to 16% total block power reduction over designs with flop trays generated by logical clustering during synthesis, and 13% clock power reduction on average compared to the previous work in [10]. Andrew B. Kahng, Jiajia Li 0002, Lutong Wang |
ICCAD | 1 |
| 2016 | BEOL stack-aware routability prediction from placement using data mining techniquesabstractIn advanced technology nodes, physical design engineers must estimate whether a standard-cell placement is routable (before invoking the router) in order to maintain acceptable design turnaround time. Modern SoC designs consume multiple compute servers, memory, tool licenses and other resources for several days to complete routing. When the design is unroutable, resources are wasted, which increases the design cost. In this work, we develop machine learning-based models that predict whether a placement solution is routable without conducting trial or early global routing. We also use our models to accurately predict iso-performance Pareto frontiers of utilization, aspect ratio and number of layers in the back-end-of-line (BEOL) stack. Furthermore, using data mining and machine learning techniques, we develop new methodologies to generate training examples given very few placements. We conduct validation experiments in three foundry technologies (28nm FDSOI, 28nm LP and 45nm GS), and demonstrate accuracy ≥ 85.9% in predicting routability of a placement. Our predictions of Pareto frontiers in the three technologies are pessimistic by at most 2% with respect to the maximum achievable utilization for a given design in a given BEOL stack. Wei-Ting Jonas Chan, Yang Du 0001, Andrew B. Kahng, Siddhartha Nath, Kambiz Samadi |
ICCD | 3 |
| 2015 | 3DIC benefit estimation and implementation guidance from 2DIC implementationabstractQuantification of three-dimensional integrated circuit (3DIC) benefits over corresponding 2DIC implementation for arbitrary designs remains a critical open problem, largely due to nonexistence of any "golden" 3DIC flow. Actual design and implementation parameters and constraints affect 2DIC and 3DIC final metrics (power, slack, etc.) in highly non-monotonic ways that are difficult for engineers to comprehend and predict. We propose a novel machine learning-based methodology to estimate 3DIC power benefit (i.e., percentage power reduction) based on corresponding golden 2DIC implementation parameters. The resulting 3D Power Estimation (3DPE) models achieve small prediction errors that are bounded by construction. We are the first to perform a novel stress test of our predictive models across a wide range of implementation and design-space parameters. Further, we explore model-guided implementation of designs in 3D to achieve minimum power: that is, our models recommend a most-promising set of implementation parameters and constraints, and also provide a priori estimates of 3D power benefits, based on a given design's post-synthesis and 2D implementation parameters. We achieve ≤10% error in power benefit prediction across various 3DIC designs. Wei-Ting Jonas Chan, Siddhartha Nath, Andrew B. Kahng, Yang Du 0001, Kambiz Samadi |
DAC | 3 |
| 2015 | Evaluation of BEOL design rule impacts using an optimal ILP-based detailed routerabstractContinued technology scaling with more pervasive use of multi-patterning has led to complex design rules and increased difficulty of maintaining high layout densities. Intuitively, emerging constraints such as unidirectional patterning or increased via spacing will decrease achievable density of the final place-and-route solution, worsening die area and product cost. However, no methodology exists for accurate assessment of design rules' impact on physical chip implementation. At the same time, this is a crucial need for early development of BEOL process technologies, particularly with FinFET or future vertical-device architectures where cell footprints can become much smaller than in bulk planar CMOS technologies. In this work, we study impacts of patterning technology choices and associated design rules on physical implementation density, with respect to cost-optimal design rule-correct detailed routing. A key contribution is an Integer Linear Programming (ILP) based optimal router (OptRouter) which considers complex design rules that arise in sub-20nm process technologies. Using OptRouter, we assess wirelength and via count impacts of various design rules (implicitly, patterning technology choices) by analyzing optimal routing solutions of clips (i.e., switchbox instances) extracted from post-detailed route layouts in an advanced technology. Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001 |
DAC | 2 |
| 2015 | A global-local optimization framework for simultaneous multi-mode multi-corner clock skew variation reductionabstractAs combinations of signoff corners grow in modern SoCs, minimization of clock skew variation across corners is important. Large skew variation can cause difficulties in multi-corner timing closure because fixing violations at one corner can lead to violations at other corners. Such "ping-pong" effects lead to significant power and area overheads and time to signoff. We propose a novel framework encompassing both global and local clock network optimizations to minimize the sum of skew variations across different PVT corners between all sequentially adjacent sink pairs. The global optimization uses linear programming to guide buffer insertion, buffer removal and routing detours. The local optimization is based on machine learning-based predictors of latency change; these are used for iterative optimization with tree surgery, buffer sizing and buffer displacement operators. Our optimization achieves up to 22% total skew variation reduction across multiple testcases implemented in foundry 28nm technology, as compared to a best-practices CTS solution using a leading commercial tool. Kwangsoo Han, Jiajia Li 0002, Andrew B. Kahng, Siddhartha Nath, Jongpil Lee |
DAC | 3 |
| 2015 | New game, new goal posts: a recent history of timing closureabstractTiming closure is the most critical phase of modern system-on-chip implementation: without timing closure, there is no tapeout. Timing closure is the end result of (i) years of methodology development, script development, signoff recipe development, etc.; (ii) months of block- and top-level final physical implementation; and (iii) a last set of manual noise and DRC fixes, with a final signoff analysis and physical verification. Over the past decade, key aspects of the underlying process and device technologies, modeling standards, EDA tooling, design methodology, and signoff criteria have changed the nature of timing closure. This paper surveys such recent evolutions in timing closure and notes directions for near-term future evolutions. Andrew B. Kahng |
DAC | 1 |
| 2015 | Optimizing Stochastic Circuits for Accuracy-Energy TradeoffsabstractStochastic computing (SC) acts on data encoded by bit-streams, and is an attractive, low-cost and error-tolerant alternative to conventional binary circuits in some important applications such as image processing and communications. We study the use of energy reduction techniques such as voltage or frequency scaling in SC circuits. We show that due to their inherent error-tolerance, SC circuits operate satisfactorily without significant accuracy loss even with aggressive scaling that improves their energy efficiency by orders of magnitude. To find the minimum-energy operating point of an SC circuit, we propose a Markov chain model that allows us to quickly explore the space of operating points. We also investigate opportunities to optimize SC circuits under such aggressive scaling. We find that logical and physical design techniques can be used to significantly expand the already powerful accuracy-energy tradeoff possibilities in SC circuits. Our simulation results show that our optimized SC circuits can tolerate aggressive voltage scaling with no significant SNR degradation after 40% supply voltage reduction (1V to 0.6V), leading to 66% energy saving (20.7pJ to 6.9pJ). Similarly, a 100% frequency boosting (400ps to 200ps) of the optimized circuits leads to no significant SNR degradation for several representative circuits. Armin Alaghi, Wei-Ting Jonas Chan, John P. Hayes, Andrew B. Kahng, Jiajia Li 0002 |
ICCAD | 4 |
| 2015 | Mixed Cell-Height Implementation for Improved Design Quality in Advanced NodesabstractIn advanced nodes, standard-cell libraries can be developed with different cell heights (e.g., in FinFET technology, corresponding to different numbers of fins). Larger cell heights provide higher drive strengths, but at the cost of larger area and power consumption as well as pin capacitance. Cells with smaller heights are relatively smaller in area, but have weaker drive strengths and are more likely to suffer from routing congestion and pin accessibility issues. Existing design methodologies and tool flows are able to mix cells with different heights at the block level (i.e., each block contains cells of a particular cell height). To our knowledge, no design methodology in the literature mixes cells of different heights in a fine-grained manner. In this work, we propose a novel physical design optimization flow to implement design blocks with mixed cell heights in a fine-grained manner. Our optimization resolves the “chicken-and-egg” loop between floorplan site definition and the optimized choices of cell heights after placement. Comprehending the constraints and costs of mixing cells of different heights (e.g., the “breaker cell” area overheads of row alignment between sub-blocks of 8T and 12T cell rows), our optimization achieves 25% area reduction versus 12T-only implementation while maintaining the same performance, and 20% performance improvement versus 8T-only implementation while maintaining similar total cell area. Sorin Dobre, Andrew B. Kahng, Jiajia Li 0002 |
ICCAD | 2 |
| 2015 | Scalable Detailed Placement Legalization for Complex Sub-14nm ConstraintsabstractTechnology scaling to 10nm and below introduces complex intra-row and inter-row constraints in standard-cell detailed placement. Examples of such constraints are found in rules for drain-drain abutment, minimum implant region area and width, oxide diffusion (OD) notching and jogging, etc. Typically, these rules are too complex for the normal global-detailed placement flow to fully consider. On the other hand, guardbanding the library cell design so that arbitrary cell placement adjacencies are all “correct by construction” has increasingly high area cost. This motivates the introduction of a final legalization phase for standard-cell placement tools in advanced (particularly 10nm and 7nm) foundry nodes. In this work, we develop a mixed integer-linear programming (MILP)-based placer, called DFPlacer, for final-phase design rule violation (DRV) fixing. DFPlacer finds (near-)DRV-free solutions considering various complex layout constraints including minimum implant width, drain-drain abutment, and oxide diffusion jogs. To overcome the runtime limitation of MILP-based approaches, we implement a distributable optimization strategy based on partitioning of the block layout into windows of cells that can be independently legalized. Using layouts in an abstracted 7nm library, we find that DFPlacer fixes 99% of DRVs on average with minimal impacts on area and timing. We also study an area-DRV tradeoff between two types of standard-cell library strategies, namely, with and without dummy poly gates. Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001 |
ICCAD | 2 |
| 2015 | Evolving EDA Beyond its E-Roots: An OverviewabstractOver the past decade, CMOS scaling has seen increasingly intrusive challenges from cost, variability, energy, reliability, and fundamental device-architectural and materials limitations. To maintain Moore's-Law scaling of integration value, the industry is urgently exploring beyond-silicon and beyond-CMOS device, interconnect and memory options, as well as heterogeneous, “More than Moore” integration and packaging technologies. This coincides with a turning point for the Electronic Design Automation (EDA) field, which has for 50+ years been a key enabler of the semiconductor industry's amazing growth. Maturation of the EDA industry and its related academic research efforts inevitably result in a spiral of declining valuations (multiples), venture capital investments, research funding, and student interest. To counter this trajectory, the EDA field's business models, research portfolios, and funding models have been going through various diversifications, but in a rather ad-hoc manner. This begs the question of how the paradigms and research methodologies of EDA can be leveraged for design automation (DA) in other, emerging domains. Arguably, more efficient evolution and growth as a community requires a more systematic, coherent effort - as well as forward-looking vision - to steer by. In this paper, we review initial efforts of a new IEEE CEDA technical activity group dedicated to this purpose. These efforts span the cataloguing of past and current research trends, development of new metrics of research impact, and visions for future applications of EDA paradigms in broader design automation contexts. It is hoped that these efforts will be useful to the EDA community as it continues to evolve beyond its “E-roots”. Andrew B. Kahng, Farinaz Koushanfar |
ICCAD | 1 |
| 2015 | Toward Metrics of Design Automation Research ImpactabstractDesign automation (DA) research has for over fifty years been performed in academia, semiconductor and system companies, and EDA companies worldwide. This research has been enabling to continued scaling of design productivity and growth of the semiconductor industry. For product companies, funding program managers and individual researchers alike, a highly relevant question is: what DA research, and what DA research outcomes, ultimately have the greatest “impact”? In this paper, we present measurements and analyses of DA research outputs (papers, patents, EDA companies), upon which future metrics of DA research impact might be based. Our studies consider 47000+ conference and journal papers from 1964-2014; the inter-patent citation graph over 759000+ DA-related patents; abstracts of 1150+ U.S. NSF projects over a three-decade span; 36 research needs documents of the Semiconductor Research Corporation from 2000-2013; and market segmentation of hundreds of EDA companies. We identify several interesting correlations, but do not claim to identify causal relationships; indeed, connecting traditional measures of research output to real-world impacts seems quite challenging. We conclude with several directions and targets for future investigation. Andrew B. Kahng, Mulong Luo, Gi-Joon Nam, Siddhartha Nath, David Z. Pan, Gabriel Robins |
ICCAD | 1 |
| 2015 | Modeling the future of semiconductors (and test!)abstractWhich semiconductor products will drive manufacturing and test technology over the next 10 to 15 years? In the past, Moore's Law has been used to predict the continuing evolution of semiconductors. Now, however, we are seeing an explosion of new device, memory and heterogeneous integration technologies aimed at achieving "More than Moore" scaling of product value. By looking at the applications that drive this explosion of new technologies, we can begin to model what the future might look like, and how the industry can deploy cost-effective manufacture and test strategies in the coming years. Three applications in particular — Smart Phone, Datacenter and IoT — will continue to have great influence on both semiconductors and systems. The ITRS "2.0" semiconductor roadmap projects these applications into the future to model potential benefits of future technologies. What will semiconductors look like in the next 10 to 15 years? Let's find out! Andrew B. Kahng |
ITC | 1 |
| 2015 | Application-Specific Cross-Layer Optimization Based on Predictive Variable-Latency VLSI DesignabstractTraditional synchronous VLSI design requires that all computations in a logic stage complete in one clock cycle. This leads to increasingly pessimistic design as technology scaling introduces increasingly significant parametric variations that result in an increasing performance variability. Alternatively, by allowing computations in a logic stage to complete in a variable number of clock cycles, variable-latency design provides relaxed timing constraints for average performance, area, and power consumption optimization. In this article, we present improved variable-latency design techniques including: (1) a generic minimum-intrusion variable-latency VLSI design paradigm, (2) a signal probability-based approximate prediction logic construction method for minimum misprediction rate at minimum cost, and (3) an application-specific cross-layer analysis methodology. Our experiments show that the proposed variable-latency design methodology on average reduces the computation latency by 26.80%(14.65%) at cost of 0.08%(3.4%) area and 0.4%(2.2%) energy consumption increase for the interger (floating point) unit of an open-source SPARC V8 processor LEON2 synthesized with a clock-cycle time between 1.97ns(3.49ns) and 5.96ns(13.74ns) based on the 45nm Nangate open cell library, while an automotive application-specific design further achieves an average latency reduction of 41.8%. Vivek De, Andrew B. Kahng, Tanay Karnik, Bao Liu 0001, Milad Maleki |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2015 | An Improved Methodology for Resilient Design ImplementationabstractResilient design techniques are used to (i) ensure correct operation under dynamic variations and to (ii) improve design performance (e.g., timing speculation). However, significant overheads (e.g., 16% and 14% energy penalties due to throughput degradation and additional circuits) are incurred by existing resilient design techniques. For instance, resilient designs require additional circuits to detect and correct timing errors. Further, when there is an error, the additional cycles needed to restore a previous correct state degrade throughput, which diminishes the performance benefit of using resilient designs. In this work, we describe an improved methodology for resilient design implementation to minimize the costs of resilience in terms of power, area, and throughput degradation. Our methodology uses two levers: selective-endpoint optimization (i.e., sensitivity-based margin insertion) and clock skew optimization. We integrate the two optimization techniques in an iterative optimization flow which comprehends toggle rate information and the trade-off between cost of resilience and margin on combinational paths. Since the error-detection network can result in up to 9% additional wirelength cost, we also propose a matching-based algorithm for construction of the error-detection network to minimize this resilience overhead. Further, our implementations comprehend the impacts of signoff corners (in particular, hold constraints, and use of typical vs. slow libraries) and process variation, which are typically omitted in previous studies of resilience trade-offs. Our proposed flow achieves energy reductions of up to 21% and 10% compared to a conventional (with only margin used to attain robustness) design and a brute-force implementation (i.e., a typical resilient design, where resilient endpoints are (greedily) instantiated at timing-critical endpoints), respectively. We show that these benefits increase in the context of an adaptive voltage scaling strategy. Andrew B. Kahng, Seokhyeong Kang, Jiajia Li 0002, José Pineda de Gyvez |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2015 | Optimization of Overdrive Signoff in High-Performance and Low-Power ICsabstractIn modern system-on-chip implementations, multimode design is commonly used to achieve better circuit performance and power across voltage-scaled, “turbo” and other operating modes. To the best of our knowledge, there is no available systematic analysis or methodology for the selection of associated signoff modes for multimode circuit implementations. In this brief, we observe significant impacts of signoff mode selection on circuit area, power, and performance. For example, incorrect choice of signoff voltages for required overdrive frequencies can incur 12% suboptimality in power or 20% in area. Using the concept of mode dominance as a guideline, we propose a scalable, model-based adaptive search methodology to explore the design space for signoff mode selection. Our proposed methodology is duty cycle-aware in its minimization of lifetime energy. Results show that our proposed methodology provides >8% improvement in performance, for given Vdd, area and power constraints, compared with the traditional “signoff and scale” method. Further, the signoff modes determined by our methods result in <;6% overhead in power compared with the optimal signoff modes. Tuck-Boon Chan, Andrew B. Kahng, Jiajia Li 0002, Siddhartha Nath, Bongil Park |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | CACTI-IO: CACTI With OFF-Chip Power-Area-Timing ModelsabstractIn this paper, we describe CACTI-IO, an extension to CACTI that includes power, area, and timing models for the IO and PHY of the OFF-chip memory interface for various server and mobile configurations. CACTI-IO enables design space exploration of the OFF-chip IO along with the dynamic random access memory and cache parameters. We describe the models added and four case studies that use CACTI-IO to study the tradeoffs between memory capacity, bandwidth (BW), and power. The case studies show that CACTI-IO helps to: 1) provide IO power numbers that can be fed into a system simulator for accurate power calculations; 2) optimize OFF-chip configurations including the bus width, number of ranks, memory data width, and OFF-chip bus frequency, especially for novel buffer-based topologies; and 3) enable architects to quickly explore new interconnect technologies, including 3-D interconnect. We find that buffers on board and 3-D technologies offer an attractive design space involving power, BW, and capacity when appropriate interconnect parameters are deployed. Norman P. Jouppi, Andrew B. Kahng, Naveen Muralimanohar, Vaishnav Srinivas |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | A deep learning methodology to proliferate golden signoff timingabstractSignoff timing analysis remains a critical element in the IC design flow. Multiple signoff corners, libraries, design methodologies, and implementation flows make timing closure very complex at advanced technology nodes. Design teams often wish to ensure that one tool's timing reports are neither optimistic nor pessimistic with respect to another tool's reports. The resulting “correlation” problem is highly complex because tools contain millions of lines of black-box and legacy code, licenses prevent any reverse-engineering of algorithms, and the nature of the problem is seemingly “unbounded” across possible designs, timing paths, and electrical parameters. In this work, we apply a “big-data” approach to the timer correlation problem. We develop a machine learning-based tool, Golden Timer eXtension (GTX), to correct divergence in flip-flop setup time, cell arc delay, wire delay, stage delay, and path slack at timing endpoints between timers. We propose a methodology to apply GTX to two arbitrary timers, and we evaluate scalability of GTX across multiple designs and foundry technologies / libraries, both with and without signal integrity analysis. Our experimental results show reduction in divergence between timing tools from 139.3ps to 21.1ps (i.e., 6.6×) in endpoint slack, and from 117ps to 23.8ps (4.9× reduction) in stage delay. We further demonstrate the incremental application of our methods so that models can be adapted to any outlier discrepancies when new designs are taped out in the same technology / library. Last, we demonstrate that GTX can also correlate timing reports between signoff and design implementation tools. Seung-Soo Han, Andrew B. Kahng, Siddhartha Nath, Ashok S. Vydyanathan |
DATE | 2 |
| 2014 | Mission profile aware IC design - A case studyabstractConsistent consideration of mission profiles throughout a supply chain is essential for the development of robust electronic components. Consideration of mission profiles is still mainly a manual task today despite rapidly decreasing robustness margins in modern automotive semiconductor technologies. Mission profile awareness aids the automation of robustness aware design by formalizing and partially automating the generation, transformation, propagation and usage of all component-specific functional loads and environmental conditions for design implementation and validation. In addition, it aids the development of electronic components in yet immature technologies or in technologies with tight parameter variation bounds. This paper introduces the general concept, requirements and context of mission profile aware design. The general design approach is presented along with key differences and enhancements to existing design approaches. A case study focusing on mission profile usage and electromigration failure avoidance is presented to demonstrate various aspects of mission profile aware design. Göran Jerke, Andrew B. Kahng |
DATE | 2 |
| 2014 | Co-optimization of memory BIST grouping, test scheduling, and logic placementabstractBuilt-in self-test (BIST) is a well-known design technique in which part of a circuit is used to test the circuit itself. BIST plays an important role for embedded memories, which do not have pins or pads exposed toward the periphery of the chip for testing with automatic test equipment. With the rapidly increasing number of embedded memories in modern SOCs (up to hundreds of memories in each hard macro of the SOC), product designers incur substantial costs of test time (subject to possible power constraints) and BIST logic physical resources (area, routing, power). However, only limited previous work addresses the physical design optimization of BIST logic; notably, Chien et al. [7] optimize BIST design with respect to test time, routing length, and area. In our work, we propose a new three-step heuristic approach to minimize test time as well as test physical layout resources, subject to given upper bounds on power consumption. A key contribution is an integer linear programming ILP framework that determines optimal test time for a given cluster of memories using either one or two BIST controllers, subject to test power limits and with full comprehension of available serialization and parallelization. Our heuristic approach integrates (i) generation of a hypergraph over the memories, with test time-aware weighting of hyperedges, along with top-down, FM-style min-cut partitioning; (ii) solution of an ILP that comprehends parallel and serial testing to optimize test scheduling per BIST controller; and (iii) placement of BIST logic to minimize routing and buffering costs. When evaluated on hard macros from a recent industrial 28nm networking SOC, our heuristic solutions reduce test time estimates by up to 11.57% with strictly fewer BIST controllers per hard macro, compared to the industrial solutions. Andrew B. Kahng, Ilgweon Kang |
DATE | 1 |
| 2014 | OCV-aware top-level clock tree optimizationabstractThe clock trees of high-performance synchronous circuits have many clock logic cells (e.g., clock gating cells, multiplexers and dividers) in order to achieve aggressive clock gating and required performance across a wide range of operating modes and conditions. As a result, clock tree structures have become very complex and difficult to optimize with automatic clock tree synthesis (CTS) tools. In advanced process nodes, CTS becomes even more challenging due to on-chip variation (OCV) effects. In this paper, we present a new CTS methodology that optimizes clock logic cell placements and buffer insertions in the top level of a clock tree. We formulate the top-level clock tree optimization problem as a linear program that minimizes a weighted sum of timing slacks, clock uncertainty and wirelength. Experimental results in a commercial 28nm FDSOI technology show that our method can improve post-CTS worst negative slack across all modes/corners by up to 320ps compared to a leading commercial provider's CTS flow. Tuck-Boon Chan, Kwangsoo Han, Andrew B. Kahng, Jae-Gon Lee, Siddhartha Nath |
ACM Great Lakes Symposium on VLSI | 3 |
| 2014 | A new methodology for reduced cost of resilienceabstractResilient design techniques are used to (i) ensure correct operation under dynamic variations; and (ii) improve design performance (e.g., through timing speculation). However, significant overheads (e.g., 17% and 15% energy penalties due to throughput degradation and additional circuits) are incurred by existing resilient design techniques. For instance, resilient designs require additional circuits to detect and correct timing errors. Further, when there is an error, the additional cycles needed to restore a previous correct state degrade throughput, which diminishes the performance benefit of using resilient designs. In this work, we propose a methodology for resilient design implementation to minimize the costs of resilience in terms of power, area and throughput degradation. Our methodology uses two levers: selective-endpoint optimization (i.e., sensitivity-based margin insertion) and clock skew optimization. We integrate the two optimization techniques in an iterative optimization flow which comprehends toggle rate information and the tradeoff between cost of resilience and margin on combinational paths. Our proposed flow achieves energy reductions of up to 19% and 21% compared to a conventional design (with only margin used to attain robustness) and a brute-force implementation, respectively. These benefits increase in the context of an adaptive voltage scaling strategy. Andrew B. Kahng, Seokhyeong Kang, Jiajia Li 0002 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2014 | Minimum implant area-aware gate sizing and placementabstractWith reduction of minimum feature size, the minimum implant area (MinIA) constraint is emerging as a new challenge for the physical implementation flow in sub-22nm technology. In particular, the MinIA constraint induces a new problem formulation wherein gate sizing and V_t-swapping must now be linked closely with detailed placement changes. To solve this new problem, we propose heuristic methods that fix MinIA violations and reduce power with gate sizing while minimizing placement perturbation to avoid creating extra timing violations. Compared to recent versions of commercial P&R tools, our methodologies achieve significant reductions (up to 100%) in the number of MinIA violations under timing/power constraints. Andrew B. Kahng, Hyein Lee 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2014 | Horizontal benchmark extension for improved assessment of physical CAD researchabstractThe rapid growth in complexity and diversity of IC designs, design flows and methodologies has resulted in a benchmark-centric culture for evaluation of performance and scalability in physicaldesign algorithm research. Landmark papers in the literature present vertical benchmarks that can be used across multiple design flow stages; artificial benchmarks with characteristics that mimic those of real designs; artificial benchmarks with known optimal solutions; as well as benchmark suites created by major companies from internal designs and/or open-source RTL. However, to our knowledge, there has been no work on horizontal benchmark creation, i.e., the creation of benchmarks that enable maximal, comprehensive assessments across commercial and academic tools at one or more specific design stages. Typically, the creation of horizontal benchmarks is limited by mismatches in data models, netlist formats, technology files, library granularity, etc. across different tools, technologies, and benchmark suites. In this paper, we describe methodology and robust infrastructure for horizontal benchmark extension" that permits maximal leverage of benchmark suites and technologies in "apples-to-apples" assessment of both industry and academic optimizers. We demonstrate horizontal benchmark extensions, and the assessments that are thus enabled, in two well-studied domains: place-and-route (four combinations of academic placers/routers, and two commercial P&R tools) and gate sizing (two academic sizers, and three commercial tools). We also point out several issues and precepts for horizontal benchmark enablement. Andrew B. Kahng, Hyein Lee 0001, Jiajia Li 0002 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2014 | Benchmarking of mask fracturing heuristicsabstractAggressive resolution enhancement techniques such as inverse lithography (ILT) often lead to complex, non-rectilinear mask shapes which make mask writing extremely slow and expensive. To reduce shot count of complex mask shapes, mask writers allow overlapping shots, due to which the problem of fracturing mask shapes with minimum shot count is NP-hard. The need to correct for e-beam proximity effect makes mask fracturing even more challenging. Although a number of fracturing heuristics have been proposed, there has been no systematic study to analyze the quality of their solutions. In this work, we propose a new method to generate benchmarks with known optimal solutions that can be used to evaluate the suboptimality of mask fracturing heuristics. We also propose a method to generate tight upper and lower bounds for actual ILT mask shapes by formulating mask fracturing as an integer linear program and solving it using branch and price. Our results show that a state-of-the-art prototype [version of] capability within a commercial EDA tool for e-beam mask shot decomposition can be suboptimal by as much as 3.7× for generated benchmarks, and by as much as 3.6× for actual ILT shapes. Tuck-Boon Chan, Puneet Gupta 0001, Kwangsoo Han, Abde Ali Kagalwalla, Andrew B. Kahng, Emile Sahouria |
ICCAD | 5 |
| 2014 | ITRS 2.0: Toward a re-framing of the Semiconductor Technology RoadmapabstractThe International Technology Roadmap for Semiconductors (ITRS) has roadmapped technology requirements of the semiconductor industry over the past two decades. The roadmap identifies major challenges in advanced technology and leads the investment of research in a cost-effective way. Traditionally, the ITRS identifies major semiconductor IC products as drivers; these set requirements for the state-of-the-art semiconductor technologies. High-performance microprocessor unit (MPU-HP) for servers and consumer portable system-on-chip (SOC-CP) for smartphones are two examples. Throughout the history of the ITRS, Moore's Law has been the main impetus for these drivers, continuously pushing the transistor density to scale at a rate of 2× per technology generation (aka “node”). However, as new requirements from applications such as data center, mobility, and context-aware computing emerge, the existing roadmapping methodology is unable to capture the entire evolution of the current semiconductor industry. Today, comprehending how key markets and applications drive the process, design and integration technology roadmap requires new system-level studies along with chip-level studies. In this paper, we extend the current ITRS roadmapping process with studies of key requirements from a system-level perspective, based on multiple generations of smartphones and microservers. We describe potential new system drivers and new metrics, and we refer to the new system-level framing of the roadmap as ITRS 2.0. Juan Antonio Carballo, Wei-Ting Jonas Chan, Paolo A. Gargini, Andrew B. Kahng, Siddhartha Nath |
ICCD | 4 |
| 2014 | Improved signoff methodology with tightened BEOL cornersabstractTo ensure functional correctness, conventional chip implementation methodology signs off the SOC design at extreme process, voltage and temperature (PVT) conditions. At the 20nm node and beyond, the back end of line (BEOL) layers have become major sources of variation, which must be accounted for by signoff at various BEOL corners. Conventional signoff methodology uses extreme BEOL corners, in which all BEOL layers are skewed to the worst-case condition (e.g., all BEOL layers have the worst parasitic capacitance). However, such a BEOL condition is very pessimistic because the probability of having all BEOL layers skew towards the worst-case condition simultaneously is extremely small. Such pessimism results in longer chip implementation schedules and poorer design quality. In this paper, we propose a signoff methodology with tightened BEOL corners to recover the pessimism incurred by the conventional BEOL corners. This approach is based on the observation that most timing-critical paths use different BEOL layers. When the variations of BEOL layers are not fully correlated, the BEOL-induced timing variation is much smaller due to averaging of random variations. Our experimental results show that by using tightened BEOL corners, we can reduce timing-violation paths by up to 100% and improve the WNS and TNS by up to 101ps and 53ns, respectively. Tuck-Boon Chan, Sorin Dobre, Andrew B. Kahng |
ICCD | 3 |
| 2014 | The ITRS MPU and SOC system drivers: Calibration and implications for design-based equivalent scaling in the roadmapabstractThe system driver models for microprocessor (MPU) and system-on-chip (SOC) in the International Technology Roadmap for Semiconductors [21] (ITRS) determine the roadmap of underlying technology requirements across devices, patterning, interconnect, test, design and other semiconductor supplier industries. In this paper, we describe several fundamental changes in the ITRS MPU and SOC system driver models as of the recently-released 2013 edition of the roadmap. We first present new A-factor (i.e., layout density) models for the logic and memory components of the MPU and SOC drivers; these updated density models comprehend the industry's shift to FinFET devices below the foundry 20nm node. We also describe updated architectural, total chip area, and total chip power models for the MPU and SOC drivers. Notably, we model the growing uncore portion of MPU products, and the growing presence of graphic processing units (GPUs) and other peripheral cores (PEs) in SOC architectures. The updated SOC architectural model enables more realistic scenario-based power modeling for the SOC driver. The 2013 ITRS update of system driver models embodies extensive calibration with foundry data as well as product structural analysis reports from a leading analysis firm (Chipworks). The model calibration reveals that the industry has contended with a “scaling gap” since 2008, whereby traditional Moore's-Law density scaling of 2× per node has failed due to patterning limitations on layout design, as well as manufacturability and performability challenges of Metal-1 half-pitch (M1HP) scaling. Growing design margins due to reliability, yield, variability, etc. have also contributed to the slowdown of density scaling. We describe how this scaling gap can potentially be compensated if the semiconductor industry urgently pursues design-based equivalent scaling (DES), which substantially changes the area and power model trajectories of MPUs and SOCs in the ITRS System Drivers Chapter. Finally, we note that as a consequence of the updated A-factor, area and power models in the 2013 ITRS, the industry now faces a 20% more daunting power management challenge than had been predicted in the 2011 roadmap. Wei-Ting Jonas Chan, Andrew B. Kahng, Siddhartha Nath, Ichiro Yamamoto |
ICCD | 2 |
| 2014 | Synthesis and Analysis of Design-Dependent Ring Oscillator (DDRO) Performance MonitorsabstractWith CMOS technology scaling, circuit performance has become more sensitive to manufacturing and environmental variations. Hence, there is a need to measure or monitor circuit performance during manufacturing and at runtime. Since each circuit may have different sensitivities to process variations, previous works have focused on the synthesis of circuit performance monitors that are specific to a given design. We develop a systematic approach for the synthesis of multiple design-dependent monitors, as well as the corresponding calibration and delay estimation methods. Our approach synthesizes design-dependent ring oscillators (DDROs) using standard-cell library gates and conventional physical implementation flows. Our delay estimation method limits the memory usage overhead by clustering critical paths with similar delay sensitivities. Experimental results show that our delay estimation method using multiple DDROs reduces overestimation (timing margin) by up to 25% compared to using a single monitor. Furthermore, our silicon measurement results for monitoring an industrial microprocessor implemented in a 45-nm silicon-on-insulator process show that DDRO can reduce the mean delay estimation error by 35% compared to inverter-based ring oscillators. Tuck-Boon Chan, Puneet Gupta 0001, Andrew B. Kahng, Liangzhen Lai |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | Optimization of overdrive signoffabstractIn modern SOC implementations, multi-mode design is commonly used to achieve better circuit performance and power across voltage-scaling, “turbo” and other operating modes. Although there are many tools for multi-mode circuit implementation, to our knowledge there is no available systematic analysis or methodology for the selection of associated signoff modes. We observe that the selection of signoff modes has significant impact on circuit area, power and performance. For example, incorrect choice of signoff voltages for required overdrive frequencies can result in a netlist with 15% suboptimality in power or 21% in area. In this paper, we propose a concept of mode dominance which can be used as a guideline for signoff mode selection. Further, we also propose efficient circuit implementation flows to optimize the selection of signoff modes within several distinct use cases. Our results show that our proposed methodology provides 5–7% improvement in performance compared to the traditional “signoff and scale” method. The signoff modes determined by our methods result in only 0.6% overhead in performance and 8% overhead in power after implementation, compared to the optimal signoff modes. Tuck-Boon Chan, Andrew B. Kahng, Jiajia Li 0002, Siddhartha Nath |
ASP-DAC | 2 |
| 2013 | On potential design impacts of electromigration awarenessabstractReliability issues significantly limit performance improvements from Moore's-Law scaling. At 45nm and below, electromigration (EM) is a serious reliability issue which affects global and local interconnects in a chip and limits performance scaling. Traditional IC implementation flows meet a 10-year lifetime requirement by overdesigning and sacrificing performance. At the same time, it is well-known among circuit designers that Black's Equation [2] suggests that lifetime can be traded for performance. In our work, we carefully study the impacts of EM-awareness on IC implementation outcomes, and show that circuit performance does not trade off so smoothly with mean time to failure (MTTF) as suggested by Black's Equation. We conduct two basic studies: EM lifetime versus performance with fixed resource budget, and EM lifetime versus resource with fixed performance. Using design examples implemented in two process nodes, we show that performance scaling achieved by reducing the EM lifetime requirement depends on the EM slack in the circuit, which in turn depends on factors such as timing constraints, length of critical paths and the mix of cell sizes. Depending on these factors, the performance gain can range from 10% to 80% when the lifetime requirement is reduced from 10 years to one year. We show that at a fixed performance requirement, power and area resources are affected by the timing slack and can either decrease by 3% or increase by 7.8% when the MTTF requirement is reduced. We also study how conventional EM fixes using per net Non-Default Rule (NDR) routing, downsizing of drivers, and fanout reduction affect performance at reduced lifetime requirements. Our study indicates, e.g., that NDR routing can increase performance by up to 5% but at the cost of 2% increase in area at a reduced 7-year lifetime requirement. Andrew B. Kahng, Siddhartha Nath, Tajana Rosing |
ASP-DAC | 1 |
| 2013 | The ITRS design technology and system drivers roadmap: process and statusabstractThe Design technology working group (TWG) is one of 16 working groups in the International Technology Roadmap for Semiconductors (ITRS) effort. It is responsible for the ITRS' Design Chapter, which roadmaps design technology requirements and potential solutions for elements of the semiconductor supply chain that are produced by the electronic design automation (EDA) industry. The Design TWG is also responsible for the ITRS' System Drivers Chapter, which roadmaps the key product classes that drive the leading-edge requirements for process and design technologies. Through these activities, the Design TWG sets a number of fundamental parameters in the overall ITRS: layout density, die size, maximum on-chip clock frequency, total chip power, SOC and MPU architecture models, etc. This paper reviews the process by which the Design TWG evolves its roadmap content, and some of the key modeling and roadmapping questions that the semiconductor and EDA industries will face in the near term. Andrew B. Kahng |
DAC | 1 |
| 2013 | Smart non-default routing for clock power reductionabstractAt advanced process nodes, non-default routing rules (NDRs) are integral to clock network synthesis methodologies. NDRs apply wider wire widths and spacings to address electromigration constraints, and to reduce parasitic and delay variations. However, wider wires result in larger driven capacitance and dynamic power. In this work, we quantify the potential for capacitance and power reduction through the application of "smart" NDR (SNDR) that substitute narrower-width NDRs on selected clock network segments, while maintaining skew, slew, delay and EM reliability criteria. We propose a practical methodology to apply smart NDRs in standard clock tree synthesis flows. Our studies with a 32/28nm library and open-source benchmarks confirm substantial (average of 9.2%) clock wire capacitance reduction and an average of 4.9% clock switching power savings over the current fixed-NDR methodology, without loss of QoR in the clock distribution. Andrew B. Kahng, Seokhyeong Kang, Hyein Lee 0001 |
DAC | 1 |
| 2013 | Impact of adaptive voltage scaling on aging-aware signoffabstractTransistor aging due to bias temperature instability (BTI) is a major reliability concern in sub-32nm technology. Aging decreases performance of digital circuits over the entire IC lifetime. To compensate for aging, designs now typically apply adaptive voltage scaling (AVS) to mitigate performance degradation by elevating supply voltage. Varying the supply voltage of a circuit using AVS also causes the BTI degradation to vary over lifetime. This presents a new challenge for margin reduction in conventional signoff methodology, which characterizes timing libraries based on transistor models with pre-calculated BTI degradations for a given IC lifetime. Many works have separately addressed predictive models of BTI and the analysis of AVS, but there is no published work that considers BTI-aware signoff that accounts for the use of AVS during IC lifetime. This motivates us to study how the presence of AVS should affect aging-aware signoff. In this paper, we first simulate and analyze circuit performance degradation due to BTI in the presence of AVS. Based on our observations, we propose a rule-of-thumb for chip designers to characterize an aging-derated standard-cell timing library that accounts for the impact of AVS. According to our experimental results, this aging-aware signoff approach avoids both overestimation and underestimation of aging - either of which results in power or area penalty - in AVS enabled systems. Tuck-Boon Chan, Wei-Ting Jonas Chan, Andrew B. Kahng |
DATE | 3 |
| 2013 | Active-mode leakage reduction with data-retained power gatingabstractPower gating is one of the most effective solutions available to reduce leakage power. However, power gating is not practically usable in an active mode due to the overheads of inrush current and data retention. In this work, we propose a data-retained power gating (DRPG) technique which enables power gating of flip-flops during active mode. More precisely, we combine clock gating and power gating techniques, with the flip-flops being power-gated during clock masked periods. We introduce a retention switch which retains data during the power gating. With the retention switch, correct logic states and functionalities are guaranteed without additional control circuitry. The proposed technique can achieve significant active-mode leakage reduction over conventional designs with small area and performance overheads. In studies with a 65nm foundry library and open-source benchmarks, DRPG achieves up to 25.7% active-mode leakage savings (11.8% savings on average) over conventional designs. Andrew B. Kahng, Seokhyeong Kang, Bongil Park |
DATE | 1 |
| 2013 | Enhanced metamodeling techniques for high-dimensional IC design estimation problemsabstractAccurate estimators of key design metrics (power, area, delay, etc.) are increasingly required to achieve IC cost reductions in system-level through physical layout optimizations. At the same time, identifying physical or analytical models of design metrics has become very challenging due to interactions among many parameters that span technology, architecture and implementation. Metamodeling techniques can simplify this problem by deriving surrogate models from samples of actual implementation data. However, the use of metamodeling techniques in IC design estimation is still in its infancy, and practitioners need more systematic understanding. In this work, we study the accuracy of metamodeling techniques across several axes: (1) low- and high-dimensional estimation problems, (2) sampling strategies, (3) sample sizes, and (4) accuracy metrics. To help obtain more general conclusions, we study these axes for three very distinct chip design estimation problems: (1) area and power of networks-on-chip routers, (2) delay and output slew of standard cells under power delivery network noise, and (3) wirelength and buffer area of clock trees. Our results show that (1) adaptive sampling can effectively reduce the sample size required to derive surrogate models by up to 64% (or, increase estimation accuracy by up to 77%) compared with Latin hypercube sampling; (2) for low-dimensional problems, Gaussian process-based models can be 1.5x more accurate than tree-based models, whereas for high-dimensional problems, tree-based models can be up to 6x more accurate than Gaussian process-based models; and (3) a variant of weighted surrogate modeling [7], which we call hybrid surrogate modeling, can improve estimation accuracy by up to 3x. Finally, to aid architects, design teams, and CAD developers in selection of the appropriate metamodeling techniques, we propose guidelines based on the insights gained from our studies. Andrew B. Kahng, Bill Lin 0001, Siddhartha Nath |
DATE | 1 |
| 2013 | High-performance gate sizing with a signoff timerabstractProcess and device scaling in late-CMOS technologies highlight leakage power as a critical challenge for the semiconductor industry. Careful gate sizing and Vth-swapping can reduce leakage, but prior optimizations based on convex or dynamic programming (i) are often based on unrealistic assumptions about circuit delay and slew propagation, (ii) fail to handle practical design rules such as transition time or load upper bounds, and (iii) do not scale well to input complexities when full extracted parasitics are available. Seeing substantial opportunities for improvement, we present a multithreaded, stochastic optimization (Trident2.0) for gate sizing and Vthassignment to minimize leakage power subject to capacitance, slew and timing constraints. Scalability and high performance of Trident2.0 are validated on ISPD-2013 Gate Sizing Contest benchmarks. Andrew B. Kahng, Seokhyeong Kang, Hyein Lee 0001, Igor L. Markov, Pankit Thapar |
ICCAD | 1 |
| 2013 | Incremental multiple-scan chain ordering for ECO flip-flop insertionabstractTestability of ECO logic is currently a significant bottleneck in the SOC implementation flow. Front-end designers sometimes require large functional ECOs close to scheduled tapeout dates or for later design revisions. To avoid loss of test coverage, ECO flip-flops must be added into existing scan chains with minimal increase to test time and minimal impact on existing routing and timing slack. We address a new Incremental Multiple-Scan Chain Ordering problem formulation to automate the tedious and time-consuming process of scan stitching for large functional ECOs. We present a heuristic with clustering, incremental clustering and ordering steps to minimize the maximum chain length (test time), routing congestion, and disturbance to existing scan chains. Test times for our incremental scan chain solutions are reduced by 5.3%, and incremental wirelength costs are reduced by 45.71%, compared to manually-solved industrial testcases. Andrew B. Kahng, Ilgweon Kang, Siddhartha Nath |
ICCAD | 1 |
| 2013 | Statistical analysis and modeling for error composition in approximate computation circuitsabstractAggressive requirements for low power and high performance in VLSI designs have led to increased interest in approximate computation. Approximate hardware modules can achieve improved energy efficiency compared to accurate hardware modules. While a number of previous works have proposed hardware modules for approximate arithmetic, these works focus on solitary approximate arithmetic operations. To utilize the benefit of approximate hardware modules, CAD tools should be able to quickly and accurately estimate the output quality of composed approximate designs. A previous work [10] proposes an interval-based approach for evaluating the output quality of certain approximate arithmetic designs. However, their approach uses sampled error distributions to store the characterization data of hardware, and its accuracy is limited by the number of intervals used during characterization. In this work, we propose an approach for output quality estimation of approximate designs that is based on a lookup table technique that characterizes the statistical properties of approximate hardwares and a regression-based technique for composing statistics to formulate output quality. These two techniques improve the speed and accuracy for several error metrics over a set of multiply-accumulator testcases. Compared to the interval-based modeling approach of [10], our approach for estimating output quality of approximate designs is 3.75ß more accurate for comparable runtime on the testcases and achieves 8.4ß runtime reduction for the error composition flow. We also demonstrate that our approach is applicable to general testcases. Wei-Ting Jonas Chan, Andrew B. Kahng, Seokhyeong Kang, Rakesh Kumar 0002, John Sartori |
ICCD | 2 |
| 2013 | Many-Core Token-Based Adaptive Power GatingabstractAmong power dissipation components, leakage power has become more dominant with each successive technology node. Leakage energy waste can be reduced by power gating. In this paper, we extend token-based adaptive power gating (TAP), a technique to power gate an actively executing core during memory accesses, to many-core Chip Multi-Processors (CMPs). TAP works by tracking every system memory request and its estimated time of arrival so that a core may power gate itself without performance or energy loss. Previous work on TAPshows several benefits compared to earlier state-of-the-art techniques, including zero performance hit and 2.58 times average energy savings for out-of-order cores. We show that TAP can adapt to increasing memory contention by increasing power-gated time by 3.69 times compared to a low memory-pressure case. We also scale TAP to many-core architectures with a distributed wake-up controller that is capable of supporting staggered wake-ups and able to power gate each core for 99.07% of the time, achieved by a non-scalable centralized scheme. Andrew B. Kahng, Seokhyeong Kang, Tajana Rosing, Richard D. Strong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2013 | Enhancing the Efficiency of Energy-Constrained DVFS DesignsabstractThe proliferation of embedded systems and mobile devices has created an increasing demand for low-energy hardware. Dynamic voltage and frequency scaling (DVFS) is a popular energy reduction technique that allows a hardware design to reduce average power consumption while still enabling the design to meet a high-performance target when necessary. To conserve energy, many DVFS-based embedded and mobile devices often spend a large fraction of their lifetimes in a low-power mode. However, DVFS designs produced by conventional multimode CAD flows tend to have significant energy overheads when operating outside of the peak performance mode, even when they are operating in a low-power mode. A dedicated core can be added for low-energy operation, but has a high cost in terms of area and leakage. In this paper, we explore the DVFS design space to identify the factors that affect DVFS efficiency. Based on our insights, we propose two design-level techniques to enhance the energy efficiency of DVFS for energy constrained systems. First, we present a context-aware DVFS design flow that considers the intrinsic characteristics of the hardware design, as well as the operating scenario—including the relative amounts of time spent in different modes, the range of performance scalability, and the target efficiency metric—to optimize the design for maximum energy efficiency. We also present a selective replication-based DVFS design methodology that identifies hardware modules for which context-aware multimode design may be inefficient and creates dedicated module replicas for different operating modes for such modules. We show that context-aware design can reduce average power by up to 20% over a conventional multimode design flow. Selective replication can reduce average power by an additional 4%. We also use the generated insights to identify microarchitectural decisions that impact DVFS efficiency. We show that the benefits from the proposed design-level techniques increase when microarchitectural transformations are allowed. Andrew B. Kahng, Seokhyeong Kang, Rakesh Kumar 0002, John Sartori |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | Accuracy-configurable adder for approximate arithmetic designsabstractApproximation can increase performance or reduce power consumption with a simplified or inaccurate circuit in application contexts where strict requirements are relaxed. For applications related to human senses, approximate arithmetic can be used to generate sufficient results rather than absolutely accurate results. Approximate design exploits a tradeoff of accuracy in computation versus performance and power. However, required accuracy varies according to applications, and 100% accurate results are still required in some situations. In this paper, we propose an accuracy-configurable approximate (ACA) adder for which the accuracy of results is configurable during runtime. Because of its configurability, the ACA adder can adaptively operate in both approximate (inaccurate) mode and accurate mode. The proposed adder can achieve significant throughput improvement and total power reduction over conventional adder designs. It can be used in accuracy-configurable applications, and improves the achievable tradeoff between performance/power and quality. The ACA adder achieves approximately 30% power reduction versus the conventional pipelined adder at the relaxed accuracy requirement. Andrew B. Kahng, Seokhyeong Kang |
DAC | 1 |
| 2012 | Explicit modeling of control and data for improved NoC router estimationabstractNetworks-on-Chip (NoCs) are scalable fabrics for interconnection networks used in many-core architectures. ORION2.0 is a widely adopted NoC power and area estimation tool; however, its models for area, power and gate count can have large errors (up to 110% on average) versus actual implementation. In this work, we propose a new methodology that analyzes netlists of NoC routers that have been placed and routed by commercial tools, and then performs explicit modeling of control and data paths followed by regression analysis to create highly accurate gate count, area and power models for NoCs. When compared with actual implementations, our new models have average estimation errors of no more than 9.8% across microarchitecture and implementation parameters. We further describe modeling extensions that enable more detailed flit-level power estimation when integrated with simulation tools such as GARNET. Andrew B. Kahng, Bill Lin 0001, Siddhartha Nath |
DAC | 1 |
| 2012 | MAPG: Memory access power gatingabstractIn mobile systems, the problems of short battery life and increased temperature are exacerbated by wasted leakage power. Leakage power waste can be reduced by power-gating a core while it is stalled waiting for a resource. In this work, we propose and model memory access power gating (MAPG), a low-overhead technique to enable power gating of an active core when it stalls during a long memory access. We describe a programmable two-stage power gating switch design that can vary a core's wake-up delay while maintaining voltage noise limits and leakage power savings. We also model the processor power distribution network and the effect of memory access power gating on neighboring cores. Last, we apply our power gating technique to actual benchmarks, and examine energy savings and overheads from power gating stalled cores during long memory accesses. Our analyses show the potential for over 38% energy savings given “perfect” power gating on memory accesses; we achieve energy savings exceeding 20% for a practical, counter-based implementation. Kwangok Jeong, Andrew B. Kahng, Seokhyeong Kang, Tajana Rosing, Richard D. Strong |
DATE | 2 |
| 2012 | Tunable sensors for process-aware voltage scalingabstractVLSI circuits usually allocate excess margin to account for worst-case process variation. Since most chips are fabricated at process conditions better than the worst-case corner, adaptive voltage scaling (AVS) is commonly used to reduce power consumption whenever possible. A typical AVS setup relies on a performance monitor that replicates critical paths of the circuit to guide voltage scaling. However, it is difficult to define appropriate critical paths for an SoC which has multiple operating modes and IPs. In this paper, we propose a different methodology for AVS which matches the voltage scaling characteristics of a circuit rather than the delays of critical paths. This fundamental change in monitoring strategy simplifies the monitoring circuitry as well as the calibration flow of conventional monitoring methods. To enable the proposed methodology, we study voltage scaling characteristics of digital circuits. Based on our analyses, we develop design guidelines as well as design monitoring circuits which have tunable voltage scaling characteristics. Our experimental results show that this methodology can be used for AVS with a simplified calibration flow. Tuck-Boon Chan, Andrew B. Kahng |
ICCAD | 2 |
| 2012 | Sensitivity-guided metaheuristics for accurate discrete gate sizingabstractThe well-studied gate-sizing optimization is a major contributor to IC power-performance tradeoffs. Viable optimizers must accurately model circuit timing, satisfy a variety of constraints, scale to large circuits, and effectively utilize a large (but finite) number of possible gate configurations, including Vt and Lg. Within the research-oriented infrastructure used in the ISPD 2012 Gate Sizing Contest, we develop a metaheuristic approach to gate sizing that integrates timing and power optimization, and handles several types of constraints. Our solutions are evaluated using a rigorous protocol that computes circuit delay with Synopsys PrimeTime. Our implementation Trident outperforms the best-reported results on all but one of the ISPD 2012 benchmarks. Compared to the 2012 contest winner, we further reduce leakage power by an average of 43%. Andrew B. Kahng, Seokhyeong Kang, Myung-Chul Kim, Igor L. Markov |
ICCAD | 2 |
| 2012 | CACTI-IO: CACTI with off-chip power-area-timing modelsabstractWe describe CACTI-IO, an extension to CACTI [4] that includes power, area and timing models for the IO and PHY of the off-chip memory interface for various server and mobile configurations. CACTI-IO enables design space exploration of the off-chip IO along with the DRAM and cache parameters. We describe the models added and three case studies that use CACTI-IO to study the tradeoffs between memory capacity, bandwidth and power. Norman P. Jouppi, Andrew B. Kahng, Naveen Muralimanohar, Vaishnav Srinivas |
ICCAD | 2 |
| 2012 | TAP: token-based adaptive power gatingabstractWe propose a low-overhead technique, Token-Based Adaptive Power Gating (TAP), to power gate an actively executing out-of-order core during memory accesses. TAP tracks every system memory request, providing a lower-bound estimate for the response time. TAP also tracks the state of every power-gateable core in the system, to provide minimal latency wake-up modes to cores such that voltage noise safety margins are not violated. A power-gating switch that utilizes TAP can deterministically power gate its core with energy savings up to 22.39% and no performance hit. Andrew B. Kahng, Seokhyeong Kang, Tajana Rosing, Richard D. Strong |
ISLPED | 1 |
| 2012 | Low-power gated bus synthesis for 3d ic via rectilinear shortest-path steiner graphabstractIn this paper, we propose a new approach for gated bus synthesis [16] with minimum wire capacitance per transaction in three-dimensional (3D) ICs. The 3D IC technology connects different device layers with through-silicon vias (TSV), which need to be considered differently from metal wire due to reliability issues and a larger footprint. Practically, the number of TSVs is bounded between layers; thus, we first devise dynamic programming and local search techniques to determine the optimal TSV locations. We then employ two approximation algorithms to generate a rectilinear shortest-path Steiner graph in each device layer. One algorithm extends the well-known greedy heuristic for the Rectilinear Steiner Arborescence problem and handles large cases with high efficiency. The other algorithm utilizes a linear programming relaxation and rounding technique which costs more time and generates a nearly-optimal Steiner graph. Experimental results show that our algorithms can construct shortest-path Steiner graphs with 22% less total wire length than the previous method of Wang et al. [16]. Chung-Kuan Cheng, Andrew B. Kahng, Shih-Hung Weng |
ISPD | 3 |
| 2012 | Construction of realistic gate sizing benchmarks with known optimal solutionsabstractGate sizing in VLSI design is a widely-used method for power or area recovery subject to timing constraints. Several previous works have proposed gate sizing heuristics for power and area optimization. However, finding the optimal gate sizing solution is NP-hard, and the suboptimality of sizing solutions has not been sufficiently quantified for each heuristic. Thus, the need for further research has been unclear. Andrew B. Kahng, Seokhyeong Kang |
ISPD | 1 |
| 2012 | Recovery-Driven Design: Exploiting Error Resilience in Design of Energy-Efficient ProcessorsabstractConventional computer-aided design (CAD) methodologies optimize a processor module for correct operation and prohibit timing violations during nominal operation. We propose recovery-driven design, a design approach that optimizes a processor module for a target timing error rate (ER) instead of correct operation. The target ER is chosen based on how many errors can be gainfully tolerated by a hardware or software error resilience mechanism. We show that significant power benefits are possible from a recovery-driven design approach that deliberately allows errors caused by voltage overscaling to occur during nominal operation, while relying on an error resilience technique to tolerate these errors. We present a detailed evaluation and analysis of such a CAD methodology that minimizes the power of a processor module for a target ER. We show how this design-level methodology can be extended to design recovery-driven processors—processors that are optimized to take advantage of hardware or software error resilience. We also discuss a gradual slack recovery-driven design approach that optimizes for a range of ERs to create soft processors—processors that have graceful failure characteristics and the ability to trade throughput or output quality for additional energy savings over a range of ERs. We demonstrate significant power benefits over conventional design—11.8% on average over all modules and ER targets, and up to 29.1% for individual modules. Processor-level benefits were 19.0%, on average. Benefits increase when recovery-driven design is coupled with an error resilience mechanism or when the number of available voltage domains increases. Andrew B. Kahng, Seokhyeong Kang, Rakesh Kumar 0002, John Sartori |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2012 | ORION 2.0: A Power-Area Simulator for Interconnection NetworksabstractAs industry moves towards multicore chips, networks-on-chip (NoCs) are emerging as the scalable fabric for interconnecting the cores. With power now the first-order design constraint, early-stage estimation of NoC power has become crucially important. In this work, we present ORION 2.0, an enhanced NoC power and area simulator, which offers significant accuracy improvement relative to its predecessor, ORION 1.0. Andrew B. Kahng, Bin Li 0018, Li-Shiuan Peh, Kambiz Samadi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | More realistic power grid verification based on hierarchical current and power constraintsabstractVectorless power grid verification algorithms, by solving linear programming (LP) problems under current constraints, enable worst-case voltage drop predictions at an early design stage. However, worst-case current patterns obtained by many existing vectorless algorithms are time-invariant (i.e., are constant throughout the simulation time), which may result in an overly pessimistic voltage drop prediction. In this paper, a more realistic power grid verification algorithm based on hierarchical current and power constraints is proposed. The proposed algorithm naturally handles general RCL power grid models. Currents at different time steps are treated as independent variables and additional power constraints are introduced; this results in more realistic time-varying worst-case current patterns and less pessimistic worst-case voltage drop predictions. Moreover, a sorting-deletion algorithm is proposed to speed up solving LP problems by utilizing the hierarchical constraint structure. Experimental results confirm that worst-case current patterns and voltage drops obtained by the proposed algorithm are more realistic, and that the sorting-deletion algorithm reduces runtime needed to solve LP problems by 85%. Chung-Kuan Cheng, Andrew B. Kahng, Grantham Pang, Yuanzhe Wang, Ngai Wong 0001 |
ISPD | 3 |
| 2010 | Slack redistribution for graceful degradation under voltage overscalingabstractModern digital IC designs have a critical operating point, or “wall of slack”, that limits voltage scaling. Even with an error-tolerance mechanism, scaling voltage below a critical voltage - so-called overscaling - results in more timing errors than can be effectively detected or corrected. This limits the effectiveness of voltage scaling in trading off system reliability and power. We propose a design-level approach to trading off reliability and voltage (power) in, e.g., microprocessor designs. We increase the range of voltage values at which the (timing) error rate is acceptable; we achieve this through techniques for power-aware slack redistribution that shift the timing slack of frequently-exercised, near-critical timing paths in a power- and area-efficient manner. The resulting designs heuristically minimize the voltage at which the maximum allowable error rate is encountered, thus minimizing power consumption for a prescribed maximum error rate and allowing the design to fail more gracefully. Compared with baseline designs, we achieve a maximum of 32.8% and an average of 12.5% power reduction at an error rate of 2%. The area overhead of our techniques, as evaluated through physical implementation (synthesis, placement and routing), is no more than 2.7%. Andrew B. Kahng, Seokhyeong Kang, Rakesh Kumar 0002, John Sartori |
ASP-DAC | 1 |
| 2010 | Improved on-chip router analytical power and area modelingabstractOver the course of this decade, uniprocessor chips have given way to multi-core chips which have become the primary building blocks of today's computer systems. The presence of multiple cores on a chip shifts the focus from computation to communication as a key bottleneck to achieving performance improvements. As industry moves towards many-core chips, networks-on-chip (NoCs) are emerging as the scalable fabric for interconnecting the cores. With power now the first-order design constraint, early-stage estimation of NoC power has become crucially important. Existing power models (e.g., ORION 2.0 [12], Xpipes [7], etc.) are based on certain router microarchitecture and circuit implementation. Therefore, when validated against different NoC prototypes - different router implementations - we saw significant deviation (up to 40% on average) that can lead to erroneous NoC design choices. This has prompted our development of a new, accurate architecture- and circuit implementation-independent router power and area modeling methodology with complete portability across existing NoC component libraries. Also, validation against a range of implemented router designs confirms substantial improvement in accuracy over existing models. Andrew B. Kahng, Bill Lin 0001, Kambiz Samadi |
ASP-DAC | 1 |
| 2010 | Eyecharts: constructive benchmarking of gate sizing heuristicsabstractDiscrete gate sizing is one of the most commonly used, flexible, and powerful techniques for digital circuit optimization. The underlying problem has been proven to be NP-hard [1]. Several (suboptimal) gate sizing heuristics have been proposed over the past two decades, but research has suffered from the lack of any systematic way of assessing the quality of the proposed algorithms. We develop a method to generate benchmark circuits (called eyecharts) of arbitrary size along with a method to compute their optimal solutions using dynamic programming. We evaluate the suboptimalities of some popular gate sizing algorithms. Eyecharts help diagnose the weaknesses of existing gate sizing algorithms, enable systematic and quantitative comparison of sizing algorithms, and catalyze further gate sizing research. Our results show that common sizing methods (including commercial tools) can be suboptimal by as much as 54% (V t -assignment), 46% (gate sizing) and 49% (gate-length biasing) for realistic libraries and circuit topologies. Puneet Gupta 0001, Andrew B. Kahng, Amarnath Kasibhatla |
DAC | 2 |
| 2010 | Recovery-driven design: a power minimization methodology for error-tolerant processor modulesabstractConventional CAD methodologies optimize a processor module for correct operation, and prohibit timing violations during nominal operation. In this paper, we propose recovery-driven design, a design approach that optimizes a processor module for a target timing error rate instead of correct operation. We show that significant power benefits are possible from a recovery-driven design flow that deliberately allows errors caused by voltage overscaling ([10],[3]) to occur during nominal operation, while relying on an error recovery technique to tolerate these errors. We present a detailed evaluation and analysis of such a CAD methodology that minimizes the power of a processor module for a target error rate. We demonstrate power benefits of up to 25%, 19%, 22%, 24%, 20%, 28%, and 20% versus traditional P&R at error rates of 0.125%, 0.25%, 0.5%, 1%, 2%, 4%, and 8%, respectively. Coupling recovery-driven design with an error recovery technique enables increased efficiency and additional power savings. Andrew B. Kahng, Seokhyeong Kang, Rakesh Kumar 0002, John Sartori |
DAC | 1 |
| 2010 | Trace-driven optimization of networks-on-chip configurationsabstractNetworks-on-chip (NoCs) are becoming increasingly important in general-purpose and application-specific multi-core designs. Although uniform router configurations are appropriate for generalpurpose NoCs, router configurations for application-specific NoCs can be non-uniformly optimized to application-specific traffic characteristics. In this paper, we specifically consider the problem of virtual channel (VC) allocation in application-specific NoCs. Prior solutions to this problem have been average-rate driven. However, average-rate models are poor representations of real application traffic, and can lead to designs that are poorly matched to the application. We propose an alternate trace-driven paradigm in which configuration of NoCs is driven by application traces. We propose two simple greedy trace-driven VC allocation schemes. Compared to uniform allocation, we observe up to 51 % reduction in the number of VCs under a given average packet latency constraint, or up to 74 % reduction in average packet latency with same number of VCs. Our results suggest that average-rate driven methods cannot effectively select appropriate links for VC allocation because they fail to consider the impact of traffic bursts. As a case study, we compare our proposed approach with an existing average-rate driven method [9] and observe up to 35 % reduction in the number of VCs for a given target latency. 1. Andrew B. Kahng, Bill Lin 0001, Kambiz Samadi, Rohit Sunkam Ramanujam |
DAC | 1 |
| 2010 | Designing a processor from the ground up to allow voltage/reliability tradeoffsabstractCurrent processor designs have a critical operating point that sets a hard limit on voltage scaling. Any scaling beyond the critical voltage results in exceeding the maximum allowable error rate, i.e., there are more timing errors than can be effectively and gainfully detected or corrected by an error-tolerance mechanism. This limits the effectiveness of voltage scaling as a knob for reliability/power tradeoffs. In this paper, we present power-aware slack redistribution, a novel design-level approach to allow voltage/reliability tradeoffs in processors. Techniques based on power-aware slack redistribution reapportion timing slack of the frequently-occurring, near-critical timing paths of a processor in a power- and area-efficient manner, such that we increase the range of voltages over which the incidence of operational (timing) errors is acceptable. This results in soft architectures — designs that fail gracefully, allowing us to perform reliability/power tradeoffs by reducing voltage up to the point that produces maximum allowable errors for our application. The goal of our optimization is to minimize the voltage at which a soft architecture encounters the maximum allowable error rate, thus maximizing the range over which voltage scaling is possible and minimizing power consumption for a given error rate. Our experiments demonstrate 23% power savings over the baseline design at an error rate of 1%. Observed power reductions are 29%, 29%, 19%, and 20% for error rates of 2%, 4%, 8%, and 16% respectively. Benefits are higher in the face of error recovery using Razor. Area overhead of our techniques is up to 2.7%. Andrew B. Kahng, Seokhyeong Kang, Rakesh Kumar 0002, John Sartori |
HPCA | 1 |
| 2010 | Efficient trace-driven metaheuristics for optimization of networks-on-chip configurationsabstractAs industry moves towards many-core chips, networks-on-chip (NoCs) are emerging as a scalable communication fabric for interconnecting the cores. With increasing core counts, there is a corresponding increase in communication demands in multi-core designs to facilitate high core utilization, and a consequent critical need for high-performance NoCs. Another megatrend in advanced technologies is that power has become the most critical design constraint. In this paper, we focus on trace-driven virtual channel (VC) allocation in application-specific NoCs. We propose a new significant VC failure metric to capture the impact of VCs on network performance and efficiently drive NoC optimization. Our proposed metaheuristics achieve up to 38% reduction in the number of VCs under a given average packet latency constraint. In addition, compared to a recently proposed trace-driven VC allocation approach, we obtain up to an O(|L|) speedup, where |L| is total number of links in the network, with no degradation in the quality of results. Andrew B. Kahng, Bill Lin 0001, Kambiz Samadi, Rohit Sunkam Ramanujam |
ICCAD | 1 |
| 2010 | Timing Yield-Aware Color Reassignment and Detailed Placement Perturbation for Bimodal CD Distribution in Double Patterning LithographyabstractDouble patterning lithography (DPL) is in current production for memory products, and is widely viewed as inevitable for logic products at the 32 nm node. DPL decomposes and prints the shapes of a critical-layer layout in two exposures. In traditional single-exposure lithography, adjacent identical layout features will have identical mean critical dimension (CD), and spatially correlated CD variations. However, with DPL, adjacent features can have distinct mean CDs, and uncorrelated CD variations. This introduces a new set of “bimodal” challenges for timing analysis and optimization. We assess the potential impact of bimodal CD distribution on timing analysis and guardbanding, and find that the traditional “unimodal” characterization and analysis framework may not be viable for DPL. We propose new bimodal-aware timing analysis and optimization methods to improve timing yield of standard-cell based designs that are manufactured using DPL. Our first contribution is a DPL-aware approach to timing modeling, based on detailed analysis of cell layouts. Our second contribution is an integer linear programming-based maximization of “alternate” mask coloring of instances in timing-critical paths, to minimize harmful covariance and performance variation. Third, we propose a dynamic programming-based detailed placement algorithm that solves mask coloring conflicts and can be used to ensure “double patterning correctness” after placement or even after detailed routing, while minimizing the displacement of timing-critical cells with manageable engineering change order (ECO) impact. With a 45 nm library and open-source design testcases, our timing-aware recoloring and placement optimization together achieve up to 271 ps (respectively, 55.75 ns) reduction in worst (respectively, total) negative slack, and 70% (respectively, 72%) reduction in worst (respectively, total) negative slack variation, respectively. Kwangok Jeong, Andrew B. Kahng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2010 | Dose Map and Placement Co-Optimization for Improved Timing Yield and Leakage PowerabstractIn sub-100nm CMOS processes, delay and leakage power reduction continue to be among the most critical design concerns. We propose to exploit the recent availability of fine-grain exposure dose control in the step-and-scan tool to achieve both design-time (placement) and manufacturing-time (yield-aware dose mapping) optimizations of timing yield and leakage power. Our placement and dose map co-optimization can improve both timing yield and leakage power of a given design. We formulate the placement-aware dose map optimization as quadratic and quadratic constraint programs which are solved using efficient quadratic program solvers. In this paper, we mainly focus on the placement-aware dose map optimization problem; in the Appendix, we describe a complementary but less impactful dose map-aware placement optimization based on an efficient cell swapping heuristic. Experimental results show noticeable improvements in minimum cycle time without leakage power increase, or in leakage power reduction without degradation of circuit performance. Kwangok Jeong, Andrew B. Kahng, Chul-Hong Park, Hailong Yao 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2010 | Layout Decomposition Approaches for Double Patterning LithographyabstractIn double patterning lithography (DPL) layout decomposition for 45 nm and below process nodes, two features must be assigned opposite colors (corresponding to different exposures) if their spacing is less than theminimum coloring spacing. However, there exist pattern configurations for which pattern features separated by less than the minimum coloring spacing cannot be assigned different colors. In such cases, DPL requires that a layout feature be split into two parts. We address this problem using two layout decomposition approaches based on aconflict graph. First, node splitting is performed at all feasible dividing points. Then, one approach detects conflict cycles in the graph which are unresolvable for DPL coloring, and determines the coloring solution for the remaining nodes using integer linear programming (ILP). The other approach, based on a different ILP problem formulation, deletes some edges in the graph to make it two-colorable, then finds the coloring solution in the new graph. We evaluate our methods on both real and artificial 45 nm testcases. Experimental results show that our proposed layout decomposition approaches effectively decompose given layouts to satisfy the key goals of minimized line-ends and maximized overlap margin. There are no design rule violations in the final decomposed layout. Andrew B. Kahng, Chul-Hong Park, Xu Xu 0001, Hailong Yao 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2010 | Accurate Predictive Interconnect Modeling for System-Level DesignabstractWe propose new accurate predictive models for the delay, power, and area of buffered interconnects to enable a more effective system-level design exploration with existing and future nanometer technology processes. We show that our models are significantly more accurate than previous models - essentially matching sign-off analyses. We integrate our models in the COSI-OCC communication synthesis infrastructure and show how they impact the feasibility and optimality of the network-on-chip architectures that are synthesized by this tool. Luca P. Carloni, Andrew B. Kahng, Sudhakar Muddu, Alessandro Pinto, Kambiz Samadi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Timing analysis and optimization implications of bimodal CD distribution in double patterning lithographyabstractDouble patterning lithography (DPL) is in current production for memory products, and is widely viewed as inevitable for logic products at the 32 nm node. DPL decomposes and prints the shapes of a critical-layer layout in two exposures. In traditional single-exposure lithography, adjacent identical layout features will have identical mean critical dimension (CD), and spatially correlated CD variations. However, with DPL, adjacent features can have distinct mean CDs, and uncorrelated CD variations. This introduces a new set of dasiabimodalpsila challenges for timing analysis and optimization. We assess the potential impact of DPL on timing analysis error and guard banding, and find that the traditional dasiaunimodalpsila characterization and analysis framework may not be viable for DPL. For example, using 45 nm models, we find that different DPL mask layout solutions can cause 50 ps skew in clock distribution that is unseen by traditional analyses. Different mask layouts can also result in 20% or more change in timing path delays. Such results lead to insights into physical design optimizations for clock and data path placement and mask coloring that can help mitigate the error and guardband costs of DPL. Kwangok Jeong, Andrew B. Kahng |
ASP-DAC | 2 |
| 2009 | ORION 2.0: A fast and accurate NoC power and area model for early-stage design space explorationabstractAs industry moves towards many-core chips, networks-on-chip (NoCs) are emerging as the scalable fabric for interconnecting the cores. With power now the first-order design constraint, early-stage estimation of NoC power has become crucially important. ORION was amongst the first NoC power models released, and has since been fairly widely used for early-stage power estimation of NoCs. However, when validated against recent NoC prototypes - the Intel 80-core Teraflops chip and the Intel scalable communications core (SCC) chip - we saw significant deviation that can lead to erroneous NoC design choices. This prompted our development of ORION 2.0, an extensive enhancement of the original ORION models which includes completely new subcomponent power models, area models, as well as improved and updated technology models. Validation against the two Intel chips confirms a substantial improvement in accuracy over the original ORION. A case study with these power models plugged within the COSI-OCC NoC design space exploration tool confirms the need for, and value of, accurate early-stage NoC power estimation. To ensure the longevity of ORION 2.0, we will be releasing it wrapped within a semi-automated flow that automatically updates its models as new technology files become available. Andrew B. Kahng, Bin Li 0018, Li-Shiuan Peh, Kambiz Samadi |
DATE | 1 |
| 2009 | Temperature- and Cost-Aware Design of 3D Multiprocessor Architecturesabstract3D stacked architectures provide significant benefits in performance, footprint and yield. However, vertical stacking increases the thermal resistances, and exacerbates temperature-induced problems that affect system reliability, performance, leakage power and cooling cost. In addition, the overhead due to through-silicon-vias (TSVs) and scribe lines contribute to the overall area, affecting wafer utilization and yield. As any of the aforementioned parameters can limit the 3D stacking process of a multiprocessor SoC (MPSoC), in this work we investigate the tradeoffs between cost and temperature profile across various technology nodes. We study how the manufacturing costs change when the number of layers, defect density, number of cores, and power consumption vary. For each design point, we also compute the steady state temperature profile, where we utilize temperature-aware floorplan optimization to eliminate the adverse effects of inefficient floorplan decisions on temperature. Our results provide guidelines for temperature-aware floorplanning in 3D MPSoCs. For each technology node, we point out the desirable design points from both cost and temperature standpoints. For example, for building a many-core SoC with 64 cores at 32 nm, stacking 2 layers provides a desirable design point. On the other hand, at 45 nm technology, stacking 3 layers keeps temperatures at an acceptable range while reducing the cost by an additional 17% in comparison to 2 layers. Ayse K. Coskun, Andrew B. Kahng, Tajana Rosing |
DSD | 2 |
| 2009 | Timing yield-aware color reassignment and detailed placement perturbation for double patterning lithographyabstractDouble patterning lithography (DPL) is a likely resolution enhancement technique for IC production in 32nm and below technology nodes. However, DPL gives rise to two independent, uncorrelated distributions of linewidth on a chip, resulting in a 'bimodal' linewidth distribution and an increase in performance variation. [13] suggested that new physical design mechanisms could reduce harmful covariance terms that contribute to this performance variation. Kwangok Jeong, Andrew B. Kahng |
ICCAD | 3 |
| 2009 | Lens aberration aware placement for timing yieldabstractProcess variations due to lens aberrations are to a large extent systematic, and can be modeled for purposes of analyses and optimizations in the design phase. Traditionally, variations induced by lens aberrations have been considered random due to their small extent. However, as process margins reduce, and as improvements in reticle enhancement techniques control variations due to other sources with increased efficacy, lens aberration-induced variations gain importance. For example, our experiments indicate that delays of most cells in the Artisan TSMC 90nm library are affected by 2--8% due to lens aberration. Aberration-induced variations are systematic and depend on the location in the lens field. In this article, we first propose an aberration-aware timing analysis flow that accounts for aberration-induced cell delay variations. We then propose an aberration-aware timing-driven analytical placement approach that utilizes the predictable slow and fast regions created on the chip due to aberration to improve cycle time. We study the dependence of our improvement on chip size, as well as use of the technique along with field blading which allows partial reticle exposure. We evaluate our technique on two testcases, AES and JPEG implemented in 90 nm technology. The proposed technique reduces cycle time by 4.322% (80ps) at the cost of 1.587% increase in trial-routed wirelength for AES. On JPEG, we observe a cycle time reduction of 5.182% (132ps) at the cost of 1.095% increase in trial-routed wirelength. Andrew B. Kahng, Chul-Hong Park, Qinke Wang |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2008 | Interconnect modeling for improved system-level design optimizationabstractAccurate modeling of delay, power, and area of interconnections early in the design phase is crucial for effective system-level optimization. Models presently used in system-level optimizations, such as network-on-chip (NoC) synthesis, are inaccurate in the presence of deep-submicron effects. In this paper, we propose new, highly accurate models for delay and power in buffered interconnects; these models are usable by system-level designers for existing and future technologies. We present a general and transferable methodology to construct our models from a wide variety of reliable sources (Liberty, LEF/ITF, ITRS, PTM, etc.). The modeling infrastructure, and a number of characterized technologies, are available as open source. Our models comprehend key interconnect circuit and layout design styles, and a power-efficient buffering technique that overcomes unrealities of previous delay-driven buffering techniques. We show that our models are significantly more accurate than previous models for global and intermediate buffered interconnects in 90 nm and 65 nm foundry processes - essentially matching signoff analyses. We also integrate our models in the COSI-OCC synthesis tool and show that the more accurate modeling significantly affects optimal/achievable architectures that are synthesized by the tool. The increased accuracy provided by our models enables system-level designers to obtain better assessments of the achievable performance/power/area tradeoffs for (communication-centric aspects of) system design, with negligible setup and overhead burdens. Luca P. Carloni, Andrew B. Kahng, Swamy Muddu, Alessandro Pinto, Kambiz Samadi |
ASP-DAC | 2 |
| 2008 | Investigation of diffusion rounding for post-lithography analysisabstractDue to aggressive scaling of device feature size to improve circuit performance in the sub-wavelength lithography regime, both diffusion and poly gate shapes are no longer rectilinear. Diffusion rounding occurs most notably where the diffusion shapes are not perfectly rectangular, including common L and T-shaped diffusion layouts to connect to power rails. This paper investigates the impact of the non-rectilinear shape of diffusion (i.e., sloped diffusion or diffusion rounding) on circuit performance (delay and leakage). Simple weighting function models for Ionmiddot and Ioffto account for the diffusion rounding effects are proposed, and compared with TCAD simulation. Our experiments show that diffusion rounding has an asymmetric characteristic for Ioff due to the differing significance of source/drain junctions on device threshold voltage. Therefore, we can model Ionmiddot and Ioffas a function of slope angle and direction. The proposed models match well with TCAD simulation results, with less than 2% and 6% error in Ionmiddot and Ioff, respectively. Puneet Gupta 0001, Andrew B. Kahng, Saumil Shah, Dennis Sylvester |
ASP-DAC | 2 |
| 2008 | Bounded-lifetime integrated circuitsabstractIntegrated circuits with bounded lifetimes can have many business advantages. We give some simple examples of methods to enforce tunable expiration dates for chips using nanometer reliability mechanisms. Puneet Gupta 0001, Andrew B. Kahng |
DAC | 2 |
| 2008 | Dose map and placement co-optimization for timing yield enhancement and leakage power reductionabstractIn sub-100nm CMOS processes, delay and leakage power reduction continue to be among the most critical design concerns. We propose to exploit the recent availability of fine-grain exposure dose control in the stepper to achieve both design-time (placement) and manufacturing-time (yield-aware dose mapping) optimizations of timing yield and leakage power. Our placement and dose map co-optimization can simultaneously improve both timing yield and leakage power of a given design. We formulate the placement-aware dose map optimization as a quadratic program, and solve it using an efficient quadratic programming solver. In this paper, we mainly focus on the placement-aware dose map optimization problem; in the Appendix, we describe the complementary but less impactful dose map-aware placement optimization, and an efficient cell swapping heuristic. Experimental results are promising: with typical 90nm stepper (ASML Dose Mapper) parameters, we achieve more than 8% improvement in minimum cycle time of the circuit without any leakage power degradation. Kwangok Jeong, Andrew B. Kahng, Chul-Hong Park, Hailong Yao 0002 |
DAC | 2 |
| 2008 | DFM in practice: hit or hype?abstractDFM has taken shape by virtue of manufacturers defining a series of "DFM activities", related to parametric and stochastic yield analysis and recommendations for design changes to improve yield. The picture is made more complex because the view from integrated device manufacturers, pure play foundries, and fabless semiconductor companies is not necessarily the same as they have different needs and constrains. This panel brings together DFM practitioners from each of these communities to discuss real experiences on the adoption level achieved so far as well as the impact in the manufactured products. The panel should be of interest of designers moving to advanced nodes (to learn from the experience of those that have "done it already") as well as those in the leading edge (such that they can compare experiences). Juan C. Rey, N. S. Nagaraj, Andrew B. Kahng, Fabian Klass, Robert C. Aitken, Cliff Hou, Luigi Capodieci |
DAC | 3 |
| 2008 | Layout decomposition for double patterning lithographyabstractIn double patterning lithography (DPL) layout decomposition for 45nm and below process nodes, two features must be assigned opposite colors (corresponding to different exposures) if their spacing is less than the minimum coloring spacing [11, 9, 5]. However, there exist pattern configurations for which pattern features separated by less than the minimum color spacing cannot be assigned different colors. In such cases, DPL requires that a layout feature be split into two parts. We address this problem using a layout decomposition algorithm that includes graph construction, conflict cycle detection, and node splitting processes. We evaluate our technique on both real-world and artificially generated testcases in 45nm technology. Experimental results show that our proposed layout decomposition method effectively decomposes given layouts to satisfy the key goals of minimized line-ends and maximized overlap margin. There are no design rule violations in the final decomposed layout. Andrew B. Kahng, Chul-Hong Park, Xu Xu 0001, Hailong Yao 0002 |
ICCAD | 1 |
| 2008 | How to get real madabstractTrue manufacturing-aware design, or "MAD", is urgently needed in light of increasing pitch, variability, leakage and reliability challenges. This talk will present directions in physical design and performance analysis that comprise "real MAD" for 45nm and below. Following are four examples. Andrew B. Kahng |
ISPD | 1 |
| 2008 | Defocus-Aware Leakage Estimation and ControlabstractLeakage power is one of the most critical issues for ultradeep submicrometer technology. Subthreshold leakage depends nearly exponentially on linewidth, and consequently, variation in linewidth translates to a large leakage variation. A significant fraction of variation in the linewidth occurs due to systematic variations involving focus and pitch. In this paper, we propose a new leakage-estimation methodology that accounts for focus-dependent variation in the linewidth. Our approach computes the pitch of each device in the design and uses it along with defocus information to predict the linewidth of the device. Once the linewidths of the devices in a cell are calculated, the cell leakage is computed to be the sum of leakages of all off-devices in the cell; device leakages are found from a linewidth-leakage table that is precharacterized with SPICE simulations. The presented methodology significantly improves the leakage estimation and can be used in existing leakage-reduction techniques to improve their efficacy. To demonstrate the use of our approach for leakage reduction, we modify the previously proposed linewidth-biasing technique of Gupta to consider the systematic variations in linewidth and further optimize the leakage power. Our method reduces the leakage spread between worst and best process corners by up to 62% compared with the conventional corner-based analysis. Defocus awareness improves the leakage reduction from linewidth-biasing by up to 7%. Andrew B. Kahng, Sudhakar Muddu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2008 | Fast Dual-Graph-Based Hotspot FilteringabstractAs advanced technologies in wafer manufacturing push patterning processes toward lower subwavelength printing, lithography for mass production potentially suffers from decreased patterning fidelity. This results in the generation of manyhotspots, which are actual device patterns with relatively large critical-dimension and image errors with respect to on-wafer targets. Hotspots can be formed under a variety of conditions such as the original design being unfriendly to the resolution enhancement technique that is applied, unanticipated pattern combinations in rule-based optical proximity correction (OPC), or inaccuracies in model-based OPC. When these hotspots fall on locations that are critical to the electrical performance of a device, device performance and parametric yield can be significantly degraded. The golden verification signoff tool using a simulation-based approach has occupied the mainstream and has been able to accurately detect hotspots. However, this approach represents a runtime-quality tradeoff point that is high in quality but also high in runtime. There is also little point in trying to replace the golden signoff tool. We are motivated to develop a low-runtime ldquoprefilterrdquo that reduces the amount of layout area to be analyzed by the golden tool, without compromising the overall quality of hotspot finding. In this paper, we first describe a novel detection algorithm for hotspots induced by lithographic uncertainty. Our goal is to rapidly detect all lithographic hotspots without significant accuracy degradation. In other words, we propose a filtering method: as long as there are no ldquofalse negatives,rdquo i.e., we reliably obtain a superset of actual hotspots, then our method can dramatically reduce the layout area processed by golden hotspot analysis. Our hotspot detection algorithm includes layout graph construction, graph planarization, three-level bridging hotspot detection, and necking hotspot detection. We have tested our flow on several industry test cases. The experimental results show that our method is promising: for benchmark designs in 90-nm and 65-nm technologies, 100% of bridging and open hotspots are detected with few falsely detected hotspots. The average runtime of our method is more than 496 faster compared to the commercial tool. Andrew B. Kahng, Chul-Hong Park, Xu Xu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2008 | CMP Fill Synthesis: A Survey of Recent StudiesabstractWe survey recent research and practice in the area of chemical-mechanical polishing (CMP) fill synthesis, in terms of both problem formulations and solution approaches. We review the CMP as the planarization technique of choice for multilevel very large-scale integration metallization processes. Post-CMP wafer topography varies according to pattern density. We review density-analysis methods and density-control objectives that are used in today's fill-synthesis algorithms. In addition, we discuss the concept of design-driven fill synthesis that seeks to optimize CMP fill with respect to objectives beyond mere density uniformity. Design-driven fill synthesis minimizes the impact of CMP fill on design performance and parametric yield while still satisfying the density criteria that are motivated by manufacturing models. We conclude with a discussion of where CMP fill synthesis may be naturally integrated within future design and manufacturing flows. Andrew B. Kahng, Kambiz Samadi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2008 | Chip Optimization Through STI-Stress-Aware Placement Perturbations and Fill InsertionabstractStarting at the 65-nm node, stress engineering to improve the performance of transistors has been a major industry focus. An intrinsic stress source-shallow trench isolation (STI)-has not been fully utilized up to now for circuit performance improvement. In this paper, we present a new methodology that combines detailed placement and active-layer fill insertion to exploit STI stress for performance improvement. We conduct process simulation of a 65-nm production STI technology to generate mobility and delay impact models for STI stress. We then utilize these models to perform STI-stress-aware delay analysis of critical paths using Simulation Program with Integrated Circuit Emphasis (SPICE). We present our timing-driven optimization of STI stress in standard cell designs, using detailed placement perturbation and active-layer fill insertion to improve complementary metal-oxide-semiconductor performance. We assess the proposed analysis and optimization on small designs implemented with a 65-nm production cell library and a standard synthesis place-and-route flow. Our stress-aware timing analysis improves the clock frequency by 4.68% to 6.31% over traditional worst case analysis, and our optimization improves clock frequency by 2.44% to 5.26%. The frequency improvement through exploitation of STI stress comes at practically zero cost in terms of design area and wire length. Andrew B. Kahng, Rasit Onur Topaloglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2007 | A Global Minimum Clock Distribution Network Augmentation Algorithm for Guaranteed Clock Skew YieldabstractNanometer VLSI systems demand robust clock distribution network design for increased process and operating condition variabilities. In this paper, we propose minimum clock distribution network augmentation for guaranteed skew yield. We present theoretical analysis results on an inserted link in a clock network, which scales down local skew and skew variation, but may not guarantee global skew and skew variation reduction in general. We propose a global minimum clock network augmentation algorithm, which inserts links simultaneously between all nearest sink pairs, apply rule-based link removal, and perform link consolidation by Steiner minimum tree construction for wirelength reduction with guaranteed clock skew yield. Our experimental results show that our proposed algorithm achieves dominant clock network augmentation solutions, e.g., an average of 16% clock skew yield improvement, 9% maximum skew reduction, and 25% reduction of clock skew variation standard deviation with identical wirelength compared with previous best clock network link insertion methods [11]. Bao Liu 0001, Andrew B. Kahng, Xu Xu 0001, Jiang Hu 0001, Ganesh Venkataraman |
ASP-DAC | 2 |
| 2007 | Line-End Shortening is Not Always a FailureabstractLine-end shortening (LES) has always been considered a catastrophic failure in circuits. However, we find that a device with some LES can continue to function correctly. Such devices have large drive current and reduced capacitance at the expense of much higher leakage current. In this paper, we investigate the power and performance characteristics of devices with LES. Our simulations show that LES does not always cause catastrophic failure of device functionality. However, in this regime LES can lead to parametric failures, aspects of which we investigate . Puneet Gupta 0001, Andrew B. Kahng, Saumil Shah, Dennis Sylvester |
DAC | 2 |
| 2007 | Design challenges at 65nm and beyond
Andrew B. Kahng |
DATE | 1 |
| 2007 | Exploiting STI stress for performanceabstractStarting at the 65 nm node, stress engineering to improve performance of transistors has been a major industry focus. An intrinsic stress source -shallow trench isolation -has not been fully utilized up to now for circuit performance improvement. In this paper, we present a new methodology that combines detailed placement and active-layer fill insertion to exploit STI stress for performance improvement. We perform process simulation of a production 65 nm STI technology to generate mobility and delay impact models for STI stress. Based on these models, we are able to perform STI stress-aware delay analysis of critical paths using SPICE. We then present our timing-driven optimization of STI stress in standard cell designs, using detailed placement perturbation to optimize PMOS performance and active-layer fill insertion to optimize NMOS performance. We assess our optimization on small designs implemented with a 65 nm production cell library and a standard synthesis, place and route flow. Our timing-driven optimization of STI stress impacts can improve clock frequency by between 7% to 11%. The frequency improvement through exploitation of STI stress comes at practically zero cost in terms of design area and wirelength. Andrew B. Kahng, Rasit Onur Topaloglu |
ICCAD | 1 |
| 2007 | Analytical thermal placement for VLSI lifetime improvement and minimum performance variationabstractDSM and nanometer VLSI designs are subject to an increasingly significant thermal effect on VLSI circuit lifetime and performance variation, which can be effectively subdued by VLSI placement. We propose analytical placement for accurate and efficient VLSI thermal optimization, and propose minimized maximum on-chip temperature as the thermal optimization objective for improved VLSI lifetime and minimized performance variation. We develop an effective analytical thermal placement technique, as well as an improved analytical placement technique with a new cell spreading function. Our experimental results show that our proposed analytical thermal placement achieves 17.85% and 30.77% maximum on-chip temperature variation reduction as well as 4.61% and 0.45% wirelength reduction respectively for the two industry design test cases compared with thermal-oblivious analytical placement, e.g., APlace. Andrew B. Kahng, Bao Liu 0001 |
ICCD | 1 |
| 2007 | Detailed placement for leakage reduction using systematic through-pitch variationabstractWe present a novel detailed placement technique that accounts for systematic through-pitch variations to reduce leakage. Leakage depends nearly exponentially on linewidth (gate length), and even small variations in linewidth introduce large variability in leakage. A substantial fraction of linewidth variation is systematic with respect to the device layout context. Detailed placement changes context of the devices that are near the cell boundaries and can be used to reduce leakage. Our approach modifies the placement of cells in small windows such that contexts that reduce leakage are created. During this optimization, cells are partitioned into rows and then placed in rows using a traveling salesman problem formulation. Andrew B. Kahng, Swamy Muddu |
ISLPED | 1 |
| 2007 | Fast and Efficient Bright-Field AAPSM Conflict Detection and CorrectionabstractAlternating-aperture phase shift masking (AAPSM), a form of strong resolution enhancement technology, will be used to image critical features on the polysilicon layer at smaller technology nodes. This technology imposes additional constraints on the layouts beyond traditional design rules. Of particular note is the requirement that all critical features be flanked by opposite-phase shifters while the shifters obey minimum width and spacing requirements. A layout is called phase assignable if it satisfies this requirement. Phase conflicts have to be removed to enable the use of AAPSM for layouts that are not phase assignable. Previous work has sought to detect a suitable set of phase conflicts to be removed as well as correct them. This paper has two key contributions: 1) a new computationally efficient approach to detect a minimal set of phase conflicts, which when corrected will produce a phase-assignable layout, and 2) a novel layout modification scheme for correcting these phase conflicts with small layout area increase. Unlike previous formulations of this problem, the proposed solution for the conflict detection problem does not frame it as a graph bipartization problem. Instead, a simpler and more computationally efficient reduction is proposed. This simplification greatly improves the runtime while maintaining the same improvements in the quality of results obtained in Chiang (Proc. DATE, 2005, p. 908). An average runtime speedup of 5.9times is achieved using the new flow. A new layout modification scheme suited for correcting phase conflicts in large standard-cell blocks is also proposed. The experiments show that the percentage area increase for making standard-cell blocks phase assignable ranges from 1.7% to 9.1% Charles C. Chiang, Andrew B. Kahng, Subarna Sinha, Xu Xu 0001, Alex Zelikovsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | Self-Compensating Design for Reduction of Timing and Leakage Sensitivity to Systematic Pattern-Dependent VariationabstractCritical dimension (CD) variation caused by defocus is largely systematic with dense lines ldquosmilingrdquo through focus while isolated lines ldquofrown.rdquo In this paper, we propose a new design methodology that allows explicit compensation of focus-dependent CD variation, in particular, either within a cell (self-compensated cells) or across cells in a critical path (self-compensated design). By creating iso and dense variants for each library cell, we can achieve designs that are more robust to focus variation. Optimization with a mixture of dense and iso cell variants is possible, both for area and leakage power in timing constraints (critical delay), with the latter an interesting complement to existing leakage-reduction techniques, such as dual-Vth. We implement both a heuristic and mixed-integer linear-programming (MILP) solution methods to address this optimization and experimentally compare their results. Results indicate that designing with a self-compensated cell library incurs 12% area penalty and 6% leakage increase over a baseline library while compensating for focus-dependent CD variation (i.e., the design meets timing constraints across a large range of focus variation). We observe 27% area penalty and 7% leakage increase at the worst case defocus condition using only single-pitch cells. The area penalty of circuits after using both the heuristic and MILP optimization approaches is reduced to 3% while maintaining timing. We also apply the optimization to leakage, which traditionally shows very large variability due to its exponential relationship with gate CD. We conclude that a mixed iso/dense library that is combined with a sensitivity-based optimization approach yields much better area/timing/leakage tradeoffs than using a self-compensated cell library alone. Self-compensated designs show 25% less leakage power on average at the worst defocus condition compared to a design employing a conventional library for the benchmarks studied. Puneet Gupta 0001, Andrew B. Kahng, Dennis Sylvester |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | Detailed Placement for Enhanced Control of Resist and Etch CDsabstractSubresolution assist feature (SRAF) and etch-dummy-insertion techniques have been absolutely essential for process-window enhancement and CD control in photo and etch processes. However, as focus levels change during lithography manufacturing, CDs at a given ldquolegalrdquo pitch can fail to achieve manufacturing tolerances. Placed standard-cell layouts may not have the ideal whitespace distribution to allow for an optimal assist-feature insertion. This paper first describes a novel dynamic-programming-based technique for Assist-Feature Correctness (AFCorr) in detailed placement of standard-cell designs. At the same time, etch-dummy features are used in the mask data preparation flow to reduce CD skew between resist and etch processes and to improve the printability of layouts. However, etch-dummy rules conflict with the SRAF insertion because each of the two techniques requires specific design rules. We further present a novel SRAF-aware etch-dummy-insertion method (SAEDM) which optimizes the etch-dummy insertion to make the layout more conducive to the assist-feature insertion after the etch-dummy features have been inserted. Since placement of cells can create forbidden-pitch violations of resist process and can increase etch skew, the placer must also generate etch-dummy-correct placement. This can be solved by Etch-dummy Correctness (EtchCorr), which is an intelligent whitespace management for etch-dummy-corrected placement, an extension of the AFCorr methodology. These methods for enhanced resist and etch CD controls are validated on industrial test cases with respect to wafer printability, database complexity, and device performance. For benchmark designs, we validate the four methodologies: 1) AFCorr; 2) SAEDM; 3) AFCorr SAEDM; and 4) AFCorr EtchCorr SAEDM. The AFCorr placement perturbation achieves a significant reduction in forbidden pitches between polysilicon shapes. Using 1) flow, forbidden-pitch count of photo process is reduced by 76%-100% for 130 nm and by 87%-100% for 90 nm. Our novel Corr design-perturbation technique, which combines the AFCorr and EtchCorr methods, facilitates additional SRAF and etch-dummy insertions and, thus, reduces the CD skew between the photo and etch processes. After Corr with SAEDM, edge-placement-error count is also reduced by 91%-100% in the resist CD and by 72%-98% in the etch CD. Our methods provide a substantial improvement in CD control with negligible timing, area, and CPU overhead. The advantages of such correctness methods are expected to increase in future technology nodes. Puneet Gupta 0001, Andrew B. Kahng, Chul-Hong Park |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | Simultaneous Buffer Insertion and Wire Sizing Considering Systematic CMP Variation and Random Leff VariationabstractAbstract—This paper presents extensions of the dynamicprogramming (DP) framework to consider buffer insertion and wire-sizing under effects of process variation. We study the effectiveness of this approach to reduce timing impact caused by chemical–mechanical planarization (CMP)-induced systematic variation and random Leffprocess variation in devices. We first present a quantitative study on the impact of CMP to interconnect parasitics. We then introduce a simple extension to handle CMP effects in the buffer insertion and wire sizing problem by simultaneously considering fill insertion (SBWF).We also tackle the same problem but with random Leffprocess variation (vSBWF) by incorporating statistical timing into the DP framework. We develop an efficient yet accurate heuristic pruning rule to approximate the computationally expensive statistical problem. Experiments under conservative assumption on process variation show that SBWF algorithm obtains 1.6% timing improvement over the variationunaware solution. Moreover, our statistical vSBWF algorithm results in 43.1% yield improvement on average. We also show that our approaches have polynomial time complexity with respect to the net-size. The proposed extensions on the DP framework is orthogonal to other power/area-constrained problems under the same framework, which has been extensively studied in the literature. Lei He 0001, Andrew B. Kahng, King Ho Tam, Jinjun Xiong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | Statistical Timing Analysis in the Presence of Signal-Integrity EffectsabstractSignal-integrity effects have significant impacts on very large-scale-integration performance variation and must be taken into account in statistical timing analysis. In this paper, we study the signal-propagation-delay variation that is induced by crosstalk aggressor signals. We establish a functional relationship between the signal propagation delay and the crosstalk aggressor signal alignment by deterministic circuit simulation and derive closed-form formulas for the statistical distributions of output signal arrival times. Our proposed method can be smoothly integrated into a static timing analyzer, wherein runtime is dominated by sampling the deterministic delay calculation, while probabilistic computation and updating take constant time. Experimental results based on the 1000- global interconnect structures in Berkeley Predictive Technology Model 70-nm technology and industry designs in 130-nm technology show that lack of statistical crosstalk aggressor signal alignment consideration could lead to up to 114.65% (71.26%) differences in interconnect-delay means (standard deviations) and 159.4% (147.4%) differences in gate-delay means (standard deviations). By contract, the method in our earlier work gives within 1.28% (3.38%) mismatch in interconnect output signal arrival time means (standard deviations) and within 2.57% (3.86%) mismatch in gate output signal arrival time means (standard deviations), respectively. Andrew B. Kahng, Bao Liu 0001, Xu Xu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2007 | Enhanced Design Flow and Optimizations for Multiproject WafersabstractThe aggressive scaling of very large-scale integration feature size and the pervasive use of advanced reticle enhancement technologies lead to dramatic increases in mask costs, pushing prototype and low-volume production designs to the limit of economic feasibility. Multiproject wafers (MPWs), or “shuttle” runs, provide an attractive solution for such designs by providing a mechanism to share the cost of mask tooling among up to tens of designs. However, MPW reticle floorplanning and wafer dicing introduce complexities that are not encountered in typical single-project wafers. Recent works on wafer dicing adopt one or more of the following assumptions to reduce problem complexity: 1) equal production volume requirement for all designs; 2) same dicing plan used for all wafers or for all rows/columns of reticle images on a wafer; 3) unrealistic wafer models such as a rectangular array of projections; and 4) fixed wafer shot-map. Although using one or more of the aforementioned assumptions makes the problem solvable, the performance of the solutions is degraded. In this paper, a comprehensive MPW flow aimed at minimizing the number of wafers needed to fulfill given die production volumes is proposed. The proposed flow includes two main steps: 1) multiproject reticle floorplanning and 2) wafer shot-map and dicing plan definition. For each of these steps, improved algorithms are proposed as follows. The proposed reticle floorplanner uses a hierarchical quadrisection combined with simulated annealing to generate “diceable” floorplans, observing given maximum reticle sizes. The proposed dicing planner allows multiple side-to-side dicing plans for different wafers and different reticle projection rows/columns within a wafer and further improves the dicing yield by partitioning each wafer into a small number of parts before individual die extraction. A wafer shot-map definition heuristic is also proposed in order to fully utilize round wafer real estate by extracting the maximum number of functional dies from both fully and partially printed reticle images. Experiments on industry test cases show that the proposed methods outperform significantly not only previous methods in the literature but also reticle floorplans manually designed by experienced engineers. Andrew B. Kahng, Ion I. Mandoiu, Xu Xu 0001, Alex Zelikovsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2007 | Stochastic Power/Ground Supply Voltage Prediction and Optimization Via Analytical PlacementabstractIncreasingly significant power/ground (P/G) supply voltage degradation in nanometer VLSI designs leads to system performance degradation and even malfunction, which requires stochastic analysis and optimization techniques. We represent the supply voltage degradation at a P/G node as a function of the supply currents and the effective resistance of a P/G supply network and propose an efficient stochastic system-level P/G supply voltage prediction method, which computes P/G supply network effective resistances in a random walk process. We further propose to reduce P/G supply voltage degradation via placement of supply current sources, and integrate P/G supply voltage degradation reduction with conventional placement objectives in an analytical placement framework. Our experimental results show that the proposed stochastic P/G supply network prediction method achieves 10x-100x speedup compared with traditional SPICE simulation, and the proposed P/G supply voltage degradation aware placement achieves an average of 20.9% (11.7%) reduction on maximum (average) supply voltage degradation with only 4.3% wirelength increase. Andrew B. Kahng, Bao Liu 0001, Qinke Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2006 | Timing-driven Steiner trees are (practically) freeabstractTraditionally, rectilinear Steiner minimum trees (RSMT) are widely used for routing estimation in design optimizations like floorplanning and physical synthesis. Since it optimizes wirelength, an RSMT may take a “non-direct ” route to a sink, which may give the designer an unnecessarily pessimistic view of the delay to the sink. Previous works have addressed this issue through performancedriven constructions, minimum Steiner arborescence, and critical sink based Steiner constructions. Physical synthesis and routing flows have been reticent to adapt universal timing-driven Steiner constructions out of fear that they are too expensive (in terms of routing resource and capacitance). This paper studies several different performance-driven Steiner tree constructions in order to show which ones have superior performance. A key result is that they add at most 2%-4 % extra capacitance, and are thus a promising avenue for today’s increasingly aggressive performance-driven P&R flows. We demonstrate using a production P&R flow that timingdriven Steiner topologies can be easily embedded into an incremental routing subflow to obtain significantly improved timing (3.6% and 5.1 % improvements in cycle time for two industry testcases) at practically no cost of wirelength or routability. Charles J. Alpert, Andrew B. Kahng, Cliff C. N. Sze, Qinke Wang |
DAC | 2 |
| 2006 | CAD challenges for leading-edge multimedia designsabstractMultimedia designs are among the most complex leading-edge integrated circuits that are made today. In this session, CAD architects for three industry-leading multimedia products will discuss new challenges that arise across the CAD flow - from system and architecture level through physical implementation - as we approach billion-transistor devices.The first talk will focus on system-level specification and codesign of embedded software for a multimedia processor. The second talk will discuss new time and space challenges for verification methodologies, which must be met to address the requirements of graphics processor time-to-market pressures. The third talk will discuss CAD challenges unique to a pioneering high-speed, multi-core processing engine for gaming and entertainment platforms. Andrew B. Kahng |
DAC | 1 |
| 2006 | DFM: where's the proof of value?abstractHow can design teams employ new tools and develop response methodologies yet still stay within design budgets? How much effort does it require to be an early adopter and what kind of measurable results compensate for this effort? Panelists discuss how their design-for-manufacture (DFM) tools fit into a fixed design methodology, budget and timeline, and give examples of expected ROI (monetary, quality, reduced time-to-market, and comprehensive yield).The aim of this panel is to provide a serious comparison of related DFM technologies on the market and some idea of the cost and difficulty of integrating the tools into a fixed design budget and timeline. Specific results will be cited, along with examples of expected ROI (monetary, quality, reduced time-to-market, and comprehensive yield enhancement).The audience should walk away with enough information to make an informed decision on which companies would make sense for their DFM challenges, to reach their own yield and throughput goals. Shishpal Rawat, Raúl Camposano, Andrew B. Kahng, Joseph Sawicki, Mike Gianfagna, Naeem Zafar, Atul Sharan |
DAC | 3 |
| 2006 | Standard cell library optimization for leakage reductionabstractScaling device geometries have caused leakage-power consumption to be one of the major challenges of deep sub-micron design and a major source for parametric yield loss. We propose a library optimization approach involving generation of additional variants for each cell master, by biasing gate-lengths of devices. We employ transistor-level gate-length assignment to exploit asymmetries in standard cell circuit topology as well slack distribution of the design. The enhanced library is used by a power optimizer to reduce design leakage without violating any timing constraints. Such transistor-level optimization of cell libraries offers significantly better leakage-delay tradeoff than simple cell-level biasing (CLB) proposed previously. Experimental results on benchmarks show transistor-level biasing (TLB) can improve the CLB leakage optimization results by 8-17%. There is a corresponding improvement in design leakage distribution as well. Saumil Shah, Puneet Gupta 0001, Andrew B. Kahng |
DAC | 3 |
| 2006 | Lens aberration aware timing-driven placementabstractProcess variations due to lens aberrations are to a large extent systematic, and can be modeled for purposes of analyses and optimizations in the design phase. Traditionally, variations induced by lens aberrations have been considered random due to their small extent. However, as process margins reduce, and as improvements in reticle enhancement techniques control variations due to other sources with increased efficacy, lens aberration-induced variations gain importance. For example, our experiments indicate that lens aberration can result in upto 8% variation in cell delay. In this paper, we propose an aberration-aware timing-driven analytical placement approach that accounts for aberration-induced variations during placement. Our approach minimizes the design's cycle time and prevents hold-time violations under systematic aberration-induced variations. On average, the proposed placement technique reduces cycle time by∼5% at the cost of∼2% increase in wirelength. Andrew B. Kahng, Chul-Hong Park, Qinke Wang |
DATE | 1 |
| 2006 | Statistical gate delay calculation with crosstalk alignment considerationabstractWe study gate delay variation caused by crosstalk aggressor alignment, i.e., difference of signal arrival times in coupled neighboring interconnects. This effect is as significant as multiple-input switching on gate delay variation [2]. We establish a functional relationship between driver gate delay and crosstalk alignment by deterministic circuit simulation, and derive closed form formulas for statistical distributions of driver gate delay and output signal arrival time.Our proposed method can be smoothly integrated into a static timing analyzer, which runtime is dominated by sampling deterministic delay calculation, while probabilistic computation and updating take constant time. Our experimental results on 70nm technology global interconnect structures and 130nm technology industry designs show respectively 159:4% and 147:4% differences in mean and standard deviation of gate delay without crosstalk aggressor alignment consideration, while our method gives within 2:57% and 3:86% offset in gate output signal arrival time mean and standard deviation, respectively. Andrew B. Kahng, Bao Liu 0001, Xu Xu 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2006 | Fill for shallow trench isolation CMPabstractShallow trench isolation (STI) is the mainstream CMOS isolation technology. It uses chemical mechanical planarization (CMP) to remove excess of deposited oxide and attain a planar surface for successive process steps. Despite advances in STI CMP technology, pattern dependencies cause large post-CMP topography variation that can result in functional and parametric yield loss. Fill insertion is used to reduce pattern variation and consequently decrease post-CMP topography variation. Traditional fill insertion is rulebased and is used with reverse etchback to attain desired planarization quality. Due to extra costs associated with reverse etchback, “single-step ” STI CMP in which fill insertion suffices is desirable. To alleviate the failures caused by imperfect CMP, we focus on two objectives for fill insertion: oxide density variation minimization and nitride density maximization. A linear programming based optimization is used to calculate oxide densities that minimize oxide density variation. Next a fill insertion methodology is presented that attains the calculated oxide density while maximizing the nitride density. Averaged over the two large testcases, the oxide density variation is reduced by 63 % and minimum nitride density is increased by 79 % compared to tiling-based fill insertion. To assess post-CMP planarization, we run CMP simulation on the layout filled with our approach and find the planarization window (time window in which polishing can be stopped) to increase by 17 % and maximum final step height (maximum difference in post-CMP oxide thickness) to decrease by 9%. Andrew B. Kahng, Alex Zelikovsky |
ICCAD | 1 |
| 2006 | Interconnect Matching Design Rule Inferring and Optimization through Correlation ExtractionabstractNew back-end design for manufacturability rules have brought guarantee rules for interconnect matching. These rules indicate a certain capacitance matching guarantee given spacing between interconnects and interconnect area. Yet, the number of these rules is so few that they are of limited value in circuit or interconnect optimization. A method to infer additional guarantees from the provided guarantees is necessary so that optimization can be optimal. In this paper, we target two problems. First, we present a methodology to infer additional matching guarantees through extracting correlation information from the given limited set of matching guarantees in the design manual. In order to achieve this, we propose a multi-function variant of multi-variate Newton-Raphson method to extract parameters of the proposed dimension-and distance-based process correlation model for interconnects. We propose to use the extracted correlation information to infer a continuum of matching rules through simulation with proposed modifications to the standard capacitance extraction procedure. Secondly, we show how to directly incorporate the inferred interconnect matching guarantees for accurate interconnect optimization in a flexible geometric programming construction. We show how much resource savings are possible through inferring of new matching rules. Applying the inferred mismatch guarantees allows a geometric programming-based H-tree optimization to reduce the clock tree resources 27% on average and up to 56%. Rasit Onur Topaloglu, Andrew B. Kahng |
ICCD | 2 |
| 2006 | Efficient decoupling capacitor planning via convex programming methodsabstractAchieving P/G supply signal integrity is crucial to success of nanometer VLSI designs. Existing P/G network optimization techniques are dominated by sensitivity based approaches. In this paper, we propose two novel convex programming based approaches for decoupling capacitor insertion in a P/G network, i.e., a semidefinite program and a linear program, which are global optimizations with theoretically guaranteed supply voltage degradation bounds. We also propose a scalability improvement scheme which enables us to apply the proposed convex programs to industry designs. We present a simple illustrative example and experimental results on an industry design, which show that the proposed semidefinite program guarantees supply voltage degradation bound for all possible supply current sources, while the proposed linear program achieves the most accurate supply voltage degradation control for a given set of supply current sources. Andrew B. Kahng, Bao Liu 0001, Sheldon X.-D. Tan |
ISPD | 1 |
| 2006 | A faster implementation of APlaceabstractAPlace is a high quality, scalable analytical placer. This paper describes our recent efforts to improve APlace for speed and scalability. We explore various wirelength and density approximation functions. We speed up the placer using a hybrid usage of wirelength and density approximaions during he course of multi-level placement, and obtain 2-2.5 imes speedup of global placement on the IBM ISPD04 and ISPD05 benchmarks. Recent applications of the APlace framework to supply voltage degradation-aware placement and lens aberration-aware timing-driven placement are also briefly described. Andrew B. Kahng, Qinke Wang |
ISPD | 1 |
| 2006 | Wafer Topography-Aware Optical Proximity CorrectionabstractDepth of focus is the major contributor to lithographic process margin. One of the major causes of focus variation is imperfect planarization of fabrication layers. Presently, optical proximity correction (OPC) methods are oblivious to the predictable nature of focus variation arising from wafer topography. As a result, designers suffer from manufacturing yield loss as well as loss of design quality through unnecessary guardbanding. In this paper, the authors propose a novel flow and method to drive OPC with a topography map of the layout that is generated by chemical-mechanical polishing simulation. The wafer topography variations result in local defocus, which the authors explicitly model in the OPC insertion and verification flows. In addition, a novel topography-aware optical rule check to validate the quality of result of OPC for a given topography is presented. The experimental validation in this paper uses simulation-based experiments with 90-nm foundry libraries and industry-strength OPC and scattering bar recipes. It is found that the proposed topography-aware OPC (TOPC) can yield up to 67% reduction in edge placement errors. TOPC achieves up to 72% reduction in worst case printability with little increase in data volume and OPC runtime. The electrical impact of the proposed TOPC method is investigated. The results show that TOPC can significantly reduce timing uncertainty in addition to process variation Puneet Gupta 0001, Andrew B. Kahng, Chul-Hong Park, Kambiz Samadi, Xu Xu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | Gate-length biasing for runtime-leakage controlabstractLeakage power has become one of the most critical design concerns for the system level chip designer. While lowered supplies (and consequently, lowered threshold voltage) and aggressive clock gating can achieve dynamic power reduction, these techniques increase the leakage power and, therefore, causes its share of total power to increase. Manufacturers face the additional challenge of leakage variability: Recent data indicate that the leakage of microprocessor chips from a single 180-nm wafer can vary by as much as 20/spl times/. Previously proposed techniques for leakage-power reduction include the use of multiple supply and gate threshold voltages, and the assignment of input values to inactive gates, such that leakage is minimized. The additional design space afforded by the biasing of device gate lengths to reduce chip leakage power and its variability is studied. It is well known that leakage power decreases exponentially and delay increases linearly with increasing gate length. Thus, it is possible to increase gate length only marginally to take advantage of the exponential leakage reduction, while impairing performance only linearly. From a design-flow standpoint, the use of only slight increases in gate length preserves both pin and layout compatibility; therefore, the authors' technique can be applied as a postlayout enhancement step. The authors apply gate-length biasing only to those devices that do not appear in critical paths, thus assuring zero or negligible degradation in chip performance. To highlight the value of the technique, the multithreshold voltage technique, which is widely used for leakage reduction, is first applied and then gate-length biasing is used to show further reduction in leakage. Experimental results show that gate-length biasing reduces leakage by 24%-38% for the most commonly used cells, while incurring delay penalties of under 10%. Selective gate-length biasing at the circuit level reduces circuit leakage by up to 30% with no delay penalty. Leakage variability is reduced significantly by up to 41%, which may lead to substantial improvements in the manufacturing yield and the product cost. The use of gate-length biasing for leakage optimization of cell instances is also assessed, in which: 1) not all timing arcs are timing critical and/or 2) the rise and fall transitions are not both timing critical at the same time. Puneet Gupta 0001, Andrew B. Kahng, Dennis Sylvester |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | Computer-Aided Optimization of DNA Array Design and ManufacturingabstractDNA probe arrays, or DNA chips, have emerged as a core genomic technology that enables cost-effective gene expression monitoring, mutation detection, single nucleotide polymorphism analysis, and other genomic analyses. DNA chips are manufactured through a highly scalable process called very large-scale immobilized polymer synthesis (VLSIPS) that combines photolithographic technologies adapted from the semiconductor industry with combinatorial chemistry. Commercially available DNA chips contain more than half a million probes and are expected to exceed 100 million probes in the next generation. This paper is one of the first attempts to apply very large scale integration (VLSI) computer-aided design methods to the physical design of DNA chips, where the main objective is to minimize total border cost (i.e., the number of nucleotide mismatches between adjacent sites). By exploiting analogies between manufacturing processes for DNA arrays and for VLSI chips, the authors demonstrate the potential for transfer of methodologies from the 40-year-old field of electronic design automation to the newer DNA array design field. The main contributions of this paper are the following. First, it proposes several partitioning-based algorithms for DNA probe placement that improve solution quality by over 4% compared to best previously known methods. Second, it gives a new design flow for DNA arrays, which enhances current methodologies by adding flow awareness to each optimization step and introducing feedback loops. Third, it proposes solution methods for new formulations integrating multiple design steps, including probe selection, placement, and embedding. Finally, it introduces new techniques to experimentally evaluate the scalability and suboptimality of existing and newly proposed probe placement algorithms. Interestingly, the authors find that DNA placement algorithms appear to have better suboptimality properties than those recently reported for VLSI placement algorithms [C.C. Chang et al., Optimality and scalability study of existing placement algorithms, Proc. Asia South-Pacific Design Automation Conf., Kitakyushu, Japan, p.621-7, Jan. 2003; J. Cong et al., Optimality, scalability and stability study of partitioning and placement algorithms, Proc. Int. Symp. Physical Design (ISPD), Monterey, CA, p.88-94, 2003] Andrew B. Kahng, Ion I. Mandoiu, Sherief Reda, Xu Xu 0001, Alex Zelikovsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2006 | New and improved BIST diagnosis methods from combinatorial Group testing theoryabstractWe examine the general problem of built-in-self-test (BIST) diagnosis in digital logic systems. The BIST diagnosis problem has applications that include identification of erroneous test vectors, faulty scan cells, and faulty items. We develop an abstract model of this problem and show a fundamental correspondence to the well-established subject of combinatorial group testing (CGT) (D. Du and F. K. Hwang, Combinatorial Group Testing and Its Applications, 1994). We exploit this new perspective to 1) link existing BIST diagnosis techniques to CGT techniques and provide further insights into existing diagnosis algorithms, 2) improve the performance of diagnosis algorithms, and 3) develop new techniques to address the BIST diagnosis problem. Using the ISCAS'89 benchmarks, we empirically demonstrate the effectiveness of our proposed techniques over existing BIST diagnosis techniques. The vastness of the CGT literature suggests that further improvements from existing research in CGT may be obtained. Andrew B. Kahng, Sherief Reda |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2006 | Wirelength minimization for min-cut placements via placement feedbackabstractThe advent of strong multilevel partitioners has made top-down min-cut placers a favored choice for modern placer implementations. Terminal propagation is an important step in min-cut placers because it translates partitioning results into global-placement wirelength assumptions. In this work, the repartitioning problem is carefully reexamined (Proc. ACM/IEEE Int. Symp. Physical Design, p. 18, 1997) in the context of terminal propagation and studied in an in-depth manner. Abstractly, it was observed that in repartitioning, future cell locations are used for present terminal propagations and that this can be conceptually regarded as a form of placement feedback. This concept was utilized to achieve accurate terminal propagation via feedback iteration and controller insertion to fine-tune the feedback response. This yields substantial reductions in placement wirelength. Implementing our approach in Capo [version 8.7 (Proc. ACM/IEEE Design Automation Conf., p. 477, 2000 and GSRC Bookshelf)] and applying it to standard benchmark circuits yields up to 14% wirelength reductions for the IBM benchmarks with an average improvement of 5.5% and up to 10% reductions for the Peko benchmarks with an average improvement of 5.37%. Experiments also show consistent improvements for routed wirelength, yielding up to 9% wirelength reductions and 5.8% average reduction with acceptable increase in placement runtime. In practice, the method proposed significantly improves routability without building congestion maps and also reduces the number of vias Andrew B. Kahng, Sherief Reda |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2006 | Zero-Change Netlist Transformations: A New Technique for Placement BenchmarkingabstractIn this paper, the authors introduce the concept of zero-change netlist transformations (ZCNTs) to: 1) quantify the suboptimality of existing placers on artificially constructed instances and 2) "partially" quantify the suboptimality of placers on synthesized netlists from arbitrary netlists by giving lower bounds to the suboptimality gap. Given a netlist and its placement from a placer, a class of netlist transformations that synthesizes a different netlist from the given netlist is formally defined, but yet the new netlist has the same half-perimeter wire length (HPWL) on the given placement. Furthermore, and more importantly, the optimal HPWL value of the new netlist is no less than that of the original netlist. By applying the transformations and reexecuting the placer, any deviation in HPWL as a lower bound to the gap from the optimal HPWL value of the new synthesized netlist can be interpreted. The transformations allow us to: 1) increase the cardinality of hyperedges; 2) reduce the number of hyperedges; and 3) increase the number of two-pin edges, while maintaining the placement HPWL constant. It is developed here methods that apply ZCNTs to synthesize netlists having typical netlist statistics. Furthermore, an approach to estimate the suboptimality of other metrics, such as rectilinear minimum-spanning tree (RMST) and minimum-Steiner tree, is extended. Using these transformations, the suboptimality of some of the existing academic placers (FengShui, Capo, mPL, Dragon) is studied on synthesized netlists from the IBM benchmarks with instances ranging from 10k to 210k placeable instances. The results show that current placers exhibit suboptimal behavior to ZCNTs with varying degree according to the placer. Systematic suboptimality deviations in HPWL and RMST are displayed on the synthesized netlists from IBM (version 1) benchmarks. The specific nature of the transformations points out troublesome netlist structures and possible directions for improvement in the existing placers Andrew B. Kahng, Sherief Reda |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2006 | A Fast Hierarchical Quadratic Placement AlgorithmabstractPlacement is a critical component of today's physical-synthesis flow with tremendous impact on the final performance of very large scale integration (VLSI) designs. Unfortunately, it accounts for a significant portion of the overall physical-synthesis runtime. With the complexity and the netlist size of today's VLSI design growing rapidly, clustering for placement can provide an attractive solution to manage affordable placement runtimes. However, such clustering has to be carefully devised to avoid any adverse impact on the final placement solution quality. This paper presents how to apply clustering and unclustering strategies to an analytic top-down placer to achieve large speedups without sacrificing (and sometimes even enhancing) the solution quality. The authors' new bottom-up clustering technique, called the best choice (BC), operates directly on a circuit hypergraph and repeatedly clusters the globally best pair of objects. Clustering score manipulation using a priority-queue (PQ) data structure enables identification of the best pair of objects whenever clustering is performed. To improve the runtime of PQ-based BC clustering, the authors proposed a lazy-update technique for faster updates of the clustering score with almost no loss of the solution quality. A number of effective methods for clustering score calculation, balancing cluster sizes, handling of fixed blocks, and area-based unclustering strategy are discussed. The effectiveness of the resulting hierarchical analytic placement algorithm is tested on several large-scale industrial benchmarks with mixed-size fixed blocks. Experimental results are promising. Compared to the flat analytic placement runs, the hierarchical mode is 2.1 times faster, on the average, with a 1.4% wire-length improvement. Gi-Joon Nam, Sherief Reda, Charles J. Alpert, Paul G. Villarrubia, Andrew B. Kahng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2005 | Panel I: who is responsible for the design for manufacturability issues in the era of nano-technologies?abstractThe notion of design for manufacturability is blurring the separation between the tasks of design and manufacture. In the era of nano-technologies, the description of the design rules has retreated back to an early stage form of many conditional cases and even an art. Thus, it is important to set the metrics for the manufacturability. However, who is going to be held accountable for the final outcomes? Should the designer, the EDA developer, the manufacture engineer, or a new breed of experts take the lead to tackle the problem? Chung-Kuan Cheng, Andrew B. Kahng, Keh-Jeng Chang, Vijay Pitchumani, Toshiyuki Shibuya, Roberto Suaya, Zhiping Yu, Fook-Luen Heng, Don MacMillen |
ASP-DAC | 3 |
| 2005 | Detailed placement for improved depth of focus and CD controlabstractSub-resolution assist features (SRAFs) provide an absolutely essential technique for critical dimension (CD) control and process window enhancement in subwavelength lithography. However, as focus levels change during manufacturing, CDs at a given "legal" pitch can fail to achieve manufacturing tolerances required for adequate yield. Furthermore, adoption of off-axis illumination (OAI) and SRAF techniques to enhance resolution at minimum pitch worsens printability of patterns at other pitches. This paper describes a novel dynamic programming-based technique for Assist-Feature Correctness (AFCorr) in detailed placement of standard-cell designs. For benchmark designs in 130nm and 90nm technologies, AFCorr achieves improved depth of focus and substantial improvement in CD control with negligible timing, area, or CPU overhead. The advantages of AFCorr are expected to increase in future technology nodes. Puneet Gupta 0001, Andrew B. Kahng, Chul-Hong Park |
ASP-DAC | 2 |
| 2005 | Power-aware placementabstractLowering power is one of the greatest challenges facing the IC industry today. We present a power-aware placement method that simultaneously performs (1) activity-based register clustering that reduces clock power by placing registers in the same leaf cluster of the clock trees in a smaller area and (2) activity-based net weighting that reduces net switching power by assigning a combination of activity and timing weights to the nets with higher switching rates or more critical timing. The method applies to designs with multiple clocks and gated clocks. We implemented the method and obtained experimental results on 8 real-world designs after placement, routing, extraction and analysis. The power-aware placement method achieved on average 25.3% and 11.4% reduction in net switching power and total power respectively, with 2.0% timing, 1.2% cell area and 11.5% runtime impact. This method has been incorporated into a commercial physical design tool. Yongseok Cheon, Pei-Hsin Ho, Andrew B. Kahng, Sherief Reda, Qinke Wang |
DAC | 3 |
| 2005 | Advanced Timing Analysis Based on Post-OPC Extraction of Critical DimensionsabstractProcess variations have become a bottleneck for predictable and high-yielding IC design and fabrication. Linewidth variation (∆L) due to defocus in a chip is largely systematic after the layout is completed, i.e., dense lines "smile" through focus while isolated (iso) lines "frown". In this paper, we propose a design flow that allows explicit compensation of focus variation, either within a cell (self-compensated cells) or across cells in a critical path (self-compensated design). Assuming that iso and dense variants are available for each library cell, we achieve designs that are more robust to focus variation. Design with a self-compensated cell library incurs ~11-12% area penalty while compensating for focus variation. Across-cell optimization with a mix of dense and iso cell variants incurs ~6-8% area overhead compared to the original cell library, while meeting timing constraints across a large range of focus variation (from 0 to 0.4um). A combination of original and iso cells provides an even better self-compensating design option, with only 1% area overhead. Circuit delay distributions are tighter with self-compensated cells and self-compensated design than with a conventional design methodology. Puneet Gupta 0001, Andrew B. Kahng, Dennis Sylvester |
DAC | 2 |
| 2005 | Bright-Field AAPSM Conflict Detection and CorrectionabstractAs feature sizes shrink, it will be necessary to use AAPSM (alternating-aperture phase shift masking) to image critical features, especially on the polysilicon layer. This imposes additional constraints on the layouts beyond traditional design rules. Of particular note is the requirement that all critical features be flanked by opposite-phase shifters, while the shifters obey minimum width and spacing requirements. A layout is called phase-assignable if it satisfies this requirement. If a layout is not phase-assignable, the phase conflicts have to be removed to enable the use of AAPSM for the layout. Previous work has sought to detect a suitable set of phase conflicts to be removed, as well as correct them. The contributions of this paper are the following: (1) a new approach to detect a minimal set of phase conflicts (also referred to as AAPSM conflicts), which when corrected will produce a phase-assignable layout; (2) a novel layout modification scheme for correcting these AAPSM conflicts. The proposed approach for conflict detection shows significant improvements in the quality of results and runtime for real industrial circuits, when compared to previous methods. To the best of our knowledge, this is the first time layout modification results are presented for bright-field AAPSM. Our experiments show that the percentage area increase for making a layout phase-assignable ranges from 0.7-11.8%. Charles C. Chiang, Andrew B. Kahng, Subarna Sinha, Xu Xu 0001, Alex Zelikovsky |
DATE | 2 |
| 2005 | Fast and efficient phase conflict detection and correction in standard-cell layoutsabstractAlternating-aperture phase shift masking (AAPSM), a form of strong resolution enhancement technology (RET) is used to image critical features on the polysilicon layer at smaller technology nodes. This technology imposes additional constraints on the layouts beyond traditional design rules. Of particular note is the requirement that all critical features be flanked by opposite-phase shifters, while the shifters obey minimum width and spacing requirements. A layout is called phase-assignable if it satisfies this requirement. Phase conflicts between shifters have to be removed to enable the use of AAPSM for layouts that air not phase-assignable. Previous work has sought to detect a suitable set of phase conflicts to be removed, as well as correct them as well as correct them. This paper has two key contributions: (1) a new computationally efficient approach to detect a minimal set of phase conflicts, which when corrected produces a phase-assignable layout; (2) a novel layout modification scheme for correcting these phase conflicts in standard-cell blocks. Unlike previous formulations of this problem, the proposed solution for the conflict detection problem does not frame it as a graph bipartization problem. Instead, a simpler and more computationally efficient reduction is proposed. This simplification greatly improves the runtime, while maintaining the same improvements ill the quality of results. An average runtime speedup of 5.9/spl times/ is achieved using the new flow. A new layout modification scheme for correcting phase conflicts in large standard-cell blocks is also proposed. The proposed layout modification scheme can handle all phase conflicts in large standard-cell blocks with small increases in area. Our experiments show that the percentage area increase for making typical standard-cell blocks phase-assignable ranges from 3.4-9.1%. Charles C. Chiang, Andrew B. Kahng, Subarna Sinha, Xu Xu 0001 |
ICCAD | 2 |
| 2005 | Intrinsic shortest path length: a new, accurate a priori wirelength estimatorabstractA priori wirelength estimation is concerned with predicting various wirelength characteristics before placement. In this work we propose a novel, accurate estimator of net lengths. We observe that in "good" placements, the length of a net is very strongly correlated with the numbers of nets in the shortest paths connecting node pairs of the net, when each shortest path is computed under the restriction that the net itself does not exist. We refer to this as the net's intrinsic shortest path length (ISPL). Using ISPL as a wirelength estimator has several advantages: (1) it transparently handles multi-pin nets and is a strong predictor of their length; (2) it strongly correlates with the average netlist wirelength; (3) it has a distribution that is similar to that of wirelength; and (4) it acts as a good predictor for individual net lengths. Based on ISPLs, we characterize VLSI netlists with a single value and develop an intuitive, empirical link between our proposed value and the Rent parameter. We also analytically model the relationship between ISPL and wirelength, and use ISPLs in two practical applications: a priori total wirelength estimation and a priori global interconnect prediction. Andrew B. Kahng, Sherief Reda |
ICCAD | 1 |
| 2005 | Architecture and details of a high quality, large-scale analytical placerabstractModern design requirements have brought additional complexities to netlists and layouts. Millions of components, whitespace resources, and fixed/movable blocks are just a few to mention in the list of complexities. With these complexities in mind, placers are faced with the burden of finding an arrangement of placeable objects under strict wirelength, timing, and power constraints. In this paper we describe the architecture and novel details of our high quality, large-scale analytical placer. The performance of our placer has been recently recognized in the recent ISPD-2005 placement contest, and in this paper we disclose many of the technical details that we believe are key factors to its performance. We describe (i) a new clustering architecture, (ii) a dynamically adaptive analytical solver, and (iii) better legalization schemes and novel detailed placement methods. We also provide extensive experimental results on a number of benchmark sets. On average, our results are better than the best published results by 3%, 14%, and 6% for the IBM ISPD '04, ICCAD '04, and ISPD '05 benchmark sets respectively. One of the goals of this paper is to also provide enough details to enable possible future replication of our methods. Andrew B. Kahng, Sherief Reda, Qinke Wang |
ICCAD | 1 |
| 2005 | Supply Voltage Degradation Aware Analytical PlacementabstractIncreasingly significant power/ground supply voltage degradation in nanometer VLSI designs leads to system performance degradation and even malfunction. Existing techniques focus on design and optimization of power/ground supply networks. In this paper, we propose supply voltage degradation aware placement, e.g., to reduce maximum supply voltage degradation by relocation of supply current sources. We represent supply voltage degradation at a P/G node as a function of supply currents and effective impedances (i.e., effective resistances in DC analysis) in a P/G network, and integrate supply voltage degradation in an analytical placement objective. For scalability and efficiency improvement, we apply random-walk, graph contraction and interpolation techniques to obtain effective resistances. Our experimental results show an average 20.9% improvement of worst-case voltage degradation and 11.7% improvement of average voltage degradation with only 4.3% wirelength increase. Andrew B. Kahng, Bao Liu 0001, Qinke Wang |
ICCD | 1 |
| 2005 | Defocus-aware leakage estimation and controlabstractLeakage power is one of the most critical issues for ultra-deep submicron technology. Subthreshold leakage depends exponentially on linewidth, and consequently variation in linewidth translates to a large leakage variation. A significant fraction of variation in linewidth occurs due to systematic variations involving focus and pitch. In this paper we propose a new leakage estimation methodology that accounts for focus-dependent variation in linewidth. The ideas presented in this paper significantly improve leakage estimation and can be used in existing leakage reduction techniques to improve their efficacy. We modify the previously proposed gate length biasing technique of [9] to consider systematic variations in linewidth and further reduce leakage power. Our method reduces the leakage spread between worst and best process corners by up to 62%. Defocus awareness improves leakage reduction from gate length biasing by up to 7% Andrew B. Kahng, Swamy Muddu |
ISLPED | 1 |
| 2005 | A semi-persistent clustering technique for VLSI circuit placementabstractPlacement is a critical component of today's physical synthesis flow with tremendous impact on the final performance of VLSI designs. However, it accounts for a significant portion of the over-all physical synthesis runtime. With complexity and netlist size of today's VLSI design growing rapidly, clustering for placement can provide an attractive solution to manage affordable placement runtime. Such clustering, however, has to be carefully devised to avoid any adverse impact on the final placement solution quality. In this paper we present a new bottom-up clustering technique, called best-choice, targeted for large-scale placement problems. Our best-choice clustering technique operates directly on a circuit hypergraph and repeatedly clusters the globally best pair of objects. Clustering score manipulation using a priority-queue data structure enables us to identify the best pair of objects whenever clustering is performed. To improve the runtime of priority-queue-based best-choice clustering, we propose a lazy-update technique for faster updates of clustering score with almost no loss of solution quality. We also discuss a number of effective methods for clustering score calculation, balancing cluster sizes, and handling of fixed blocks. The effectiveness of our best-choice clustering methodology is demonstrated by extensive comparisons against other standard clustering techniques such as Edge-Coarsening [12] and First-Choice [13]. All clustering methods are implemented within an industrial placer CPLACE [1] and tested on several industrial benchmarks in a semi-persistent clustering context. Charles J. Alpert, Andrew B. Kahng, Gi-Joon Nam, Sherief Reda, Paul G. Villarrubia |
ISPD | 2 |
| 2005 | Simultaneous buffer insertion and wire sizing considering systematic CMP variation and random leff variationabstractAbstract—This paper presents extensions of the dynamic-programming (DP) framework to consider buffer insertion and wire-sizing under effects of process variation. We study the effectiveness of this approach to reduce timing impact caused by chemical–mechanical planarization (CMP)-induced systematic variation and random Leff process variation in devices. We first present a quantitative study on the impact of CMP to interconnect parasitics. We then introduce a simple extension to handle CMP effects in the buffer insertion and wire sizing problem by simulta-neously considering fill insertion (SBWF). We also tackle the same problem but with random Leff process variation (vSBWF) by in-corporating statistical timing into the DP framework. We develop an efficient yet accurate heuristic pruning rule to approximate the computationally expensive statistical problem. Experiments under conservative assumption on process variation show that SBWF algorithm obtains 1.6 % timing improvement over the variation-unaware solution. Moreover, our statistical vSBWF algorithm results in 43.1 % yield improvement on average. We also show that our approaches have polynomial time complexity with respect to the net-size. The proposed extensions on the DP framework is orthogonal to other power/area-constrained problems under the same framework, which has been extensively studied in the literature. Index Terms—Buffering, dummy fill insertion, fill patterns, interconnect optimization, process variation, random Leff variation, systematic CMP variation, wire sizing. I. Lei He 0001, Andrew B. Kahng, King Ho Tam, Jinjun Xiong |
ISPD | 2 |
| 2005 | Evaluation of placer suboptimality via zero-change netlist transformationsabstractIn this paper we introduce the concept of zero-change transformations to quantify the suboptimality of existing placers. Given a netlist and its placement from a placer, we formally define a class of netlist transformations that produce different netlists from the given netlist but have the same Half-Perimeter Wire Length (HPWL). Furthermore, the optimal HPWL value of the new netlists is no less than that of the original netlist. By applying our transformations and re-executing the placer, we can interpret any deviation in HPWL as a lower bound to the deviation from the optimal HPWL value. Such deviation is a measure of suboptimality. Using these transformations, the suboptimality of several existing academic and industrial placers is studied on the IBM benchmarks. Our results show that current placers are sub-optimal for zero-change transformations with deviations in HPWL by up to 32% on the IBM (version 1) benchmarks. The specific nature of our transformations also pinpoints possible directions for improvement in existing placers. Andrew B. Kahng, Sherief Reda |
ISPD | 1 |
| 2005 | APlace: a general analytic placement frameworkabstractWe streamline and extend APlace, the general analytic placement engine based on ideas of Naylor et al. [7] and described in [3, 4, 5]. Previous work explored the adaptability of APlace to multiple contexts with good quality of results. For example, the framework was extended to traditional wirelength-driven standard-cell placement in [3, 5], achieving good results in placed HPWL and routed final wire-length. The framework was also extended to top-down multilevel placement, congestion-directed placement, mixed-size placement, timing-driven placement, I/O-core co-placement and constraint handling for mixed-signal contexts [3, 4, 5]. In this work, we have modified the implementation of APlace for speed and scalability. Improvements have been made in clustering, legalization and detailed placement strategies, as well as via a distributable solution framework for both global and detailed placement phases. Andrew B. Kahng, Sherief Reda, Qinke Wang |
ISPD | 1 |
| 2005 | The Y architecture for on-chip interconnect: analysis and methodologyabstractThe Y architecture for on-chip interconnect is based on pervasive use of 0/spl deg/, 120/spl deg/, and 240/spl deg/ oriented semiglobal and global wiring. Its use of three uniform directions exploits on-chip routing resources more efficiently than traditional Manhattan wiring architecture. This paper gives in-depth analysis of deployment issues associated with the Y architecture. Our contributions are as follows. 1) We analyze communication capability (throughput of meshes) for different interconnect architectures using a multicommodity flow approach and a Rentian communication model. Throughput of the Y architecture is largely improved compared to the Manhattan architecture, and is close to the throughput of the X architecture. 2) We improve existing estimates for the wirelength reduction of various interconnect architectures by taking into account the effect of routing-geometry-aware placement. 3) We propose a symmetrical Y clock tree structure with better total wire length compared to both H and X clock tree structures, and better path length compared to the H tree. 4) We discuss power distribution under the Y architecture, and give analytical and SPICE simulation results showing that the power network in Y architecture can achieve (8.5%) less IR drop than an equally resourced power network in Manhattan architecture. 5) We propose the use of via tunnels and banks of via tunnels as a technique for improving routability for Manhattan and Y architectures. Hongyu Chen 0001, Chung-Kuan Cheng, Andrew B. Kahng, Ion I. Mandoiu, Qinke Wang, Bo Yao 0004 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2005 | Compressible area fill synthesisabstractControl of variability and performance in the back end of the VLSI manufacturing line has become extremely difficult with the introduction of new materials such as copper and low-k dielectrics. To improve manufacturability, and in particular to enable more uniform chemical-mechanical planarization (CMP), it is necessary to insert area fill features into low-density layout regions. Because area fill feature sizes are very small compared to the large empty layout areas that need to be filled, the filling process can increase the size of the resulting layout data file by an order of magnitude or more. To reduce file transfer times, and to accommodate future maskless lithography regimes, data compression becomes a significant requirement for fill synthesis. In this paper, we make the following contributions. First, we define two complementary strategies for fill data volume reduction corresponding to two different points in the design-to-manufacturing flow: compressible filling and post-fill compression . Second, we compare compressible filling methods in the fixed-dissection regime when two different sets of compression operators are used: the traditional GDSII array reference (AREF) construct, and the new Open Artwork System Interchange Standard (OASIS) repetitions. We apply greedy techniques to find practical compressible filling solutions and compare them with optimal integer linear programming solutions. Third, for the post-fill data compression problem, we propose two greedy heuristics, an exhaustive search-based method, and a smart spatial regularity search technique. We utilize an optimal bipartite matching algorithm to apply OASIS repetition operators to irregular fill patterns. Our experimental results indicate that both fill data compression methodologies can achieve significant data compression ratios, and that they outperform industry tools such as Calibre V8.8 from Mentor Graphics. Our experiments also highlight the advantages of the new OASIS compression operators over the GDSII AREF construct. Yu Chen 0005, Andrew B. Kahng, Gabriel Robins, Alex Zelikovsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2005 | Layout-aware scan chain synthesis for improved path delay fault coverageabstractPath delay fault testing has become increasingly important due to higher clock rates and higher process variability caused by shrinking geometries. Achieving high-coverage path delay fault testing requires the application of scan justified test vector pairs, coupled with careful ordering of the scan flip-flops and/or insertion of dummy flip-flops in the scan chain. Previous works on scan synthesis for path delay fault testing using scan shifting have focused exclusively on maximizing fault coverage and/or minimizing the number of dummy flip-flops, but have disregarded the scan wirelength overhead. In this paper we propose a layout-aware coverage-driven scan chain ordering methodology and give exact and heuristic algorithms for computing the achievable tradeoffs between path delay fault coverage and both dummy flip-flop and wirelength costs. Experimental results show that our scan chain ordering methodology yields significant improvements in path delay coverage with a very small increase in wirelength overhead compared to previous layout-driven approaches, and similar coverage with up to 25 times improvement in wirelength compared to previous layout-oblivious coverage-driven approaches. Puneet Gupta 0001, Andrew B. Kahng, Ion I. Mandoiu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2005 | Implementation and extensibility of an analytic placerabstractAutomated cell placement is a critical problem in very large scale integration (VLSI) physical design. New analytical placement methods that simultaneously spread cells and optimize wirelength have recently received much attention from both academia and industry. A novel and simple objective function for spreading cells over the placement area is described in the patent of Naylor et al. (U.S. Pat. 6301693). When combined with a wirelength objective function, this allows efficient simultaneous cell spreading and wirelength optimization using nonlinear optimization techniques. In this paper, we implement an analytic placer (APlace) according to these ideas (which have other precedents in the open literature), and conduct in-depth analysis of characteristics and extensibility of the placer. Our contributions are as follows. 1) We extend the objective functions described in (Naylor et al., U.S. Patent 6301693) with congestion information and implement a top-down hierarchical (multilevel) placer (APlace) based on them. For IBM-ISPD04 circuits, the half-perimeter wirelength of APlace outperforms that of FastPlace, Dragon, and Capo, respectively, by 7.8%, 6.5%, and 7.0% on average. For eight IBM-PLACE v2 circuits, after the placements are detail-routed using Cadence WRoute, the average improvement in final wirelength is 12.0%, 8.1%, and 14.1% over QPlace, Dragon, and Capo, respectively. 2) We extend the placer to address mixed-size placement and achieve an average of 4% wirelength reduction on ten ISPD'02 mixed-size benchmarks compared to results of the leading-edge solver, FengShui. 3) We extend the placer to perform timing-driven placement. Compared with timing-driven industry tools, evaluated by commercial detailed routing and static timing analysis, we achieve an average of 8.4% reduction in cycle time and 7.5% reduction in wirelength for a set of six industry testcases. 4) We also extend the placer to perform input/output-core coplacement and constraint handing for mixed-signal designs. Our paper aims to, and empirically demonstrates, that the APlace framework is a general, and extensible platform for "spatial embedding" tasks across many aspects of system physical implementation. Andrew B. Kahng, Qinke Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | Routing-aware scan chain orderingabstractScan chain insertion can have a large impact on routability, wirelength, and timing of the design. We present a routing-driven methodology for scan chain ordering with minimum wirelength objective. A routing-based approach to scan chain ordering, while potentially more accurate, can result in TSP (Traveling Salesman Problem) instances which are asymmetric and highly nonmetric; this may require a careful choice of solvers. We evaluate our new methodology on recent industry place-and-route blocks with 1200 to 5000 scan cells. We show substantial wirelength reductions for the routing-based flow versus the traditional placement-based flow. In a number of our test cases, over 86% of scan routing overhead is saved. Even though our experiments are, so far, timing oblivious, the routing-based flow also improves evaluated timing, and practical timing-driven extensions appear feasible. Puneet Gupta 0001, Andrew B. Kahng, Stefanus Mantik |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2004 | Optimal planning for mesh-based power distribution
Hongyu Chen 0001, Chung-Kuan Cheng, Andrew B. Kahng, Makoto Mori, Qinke Wang |
ASP-DAC | 3 |
| 2004 | Combinatorial group testing methods for the BIST diagnosis problem
Andrew B. Kahng, Sherief Reda |
ASP-DAC | 1 |
| 2004 | Quantum-Dot Cellular Automata (QCA) circuit partitioning: problem modeling and solutionsabstractThis paper presents the Quantum-Dot Cellular Automata (QCA) physical design problem, in the context of the VLSI physical design problem. The problem is divided into three subproblems: partitioning, placement, and routing of QCA circuits. This paper presents an ILP formulation and heuristic solution to the partitioning problem, and compares the two sets of results. Additionally, we compare a human-generated circuit to the ILP and Heuristic solutions. The results demonstrate that the heuristic is a practical method of reducing partitioning run time while providing a result that is close to the optimal for a given circuit. Dominic A. Antonelli, Danny Ziyi Chen, Timothy J. Dysart, Xiaobo Sharon Hu, Andrew B. Kahng, Peter M. Kogge, Richard C. Murphy, Michael T. Niemier |
DAC | 5 |
| 2004 | Toward a methodology for manufacturability-driven design rule explorationabstractResolution enhancement techniques (RET) such as optical proximity correction (OPC) and phase-shift mask (PSM) technology are deployed in modern processes to increase the fidelity of printed features, especially critical dimensions (CD) in polysilicon. Even given these exotic technologies, there has been momentum towards less exibility in layout, in order to ensure printability. However, there has not been a systematic study of the performance and manufacturability impact of such a move towards restrictive design rules. In this paper we present a design ow that evaluates the application of various restricted design rule (RDR) sets in deep submicron ASIC designs in terms of circuit performance and parametric yield. Using such a framework, process and design engineers can identify potential solutions to maximize manufacturability by selectively applying RDRs while maintaining chip performance. In this work we focus attention on the device layer which is the most difficult design layer to manufacture. We quantify the performance, manufacturability and mask cost impact of several common design rules. For instance, we find that small increases in the minimum allowable poly line end extension beyond active provide high levels of immunity to lithographic defocus conditions. Also, modification of the minimum field poly to diffusion spacing can provide good manufacturability, while a single pitch single orientation design rule can reduce gate 3σ uncertainty. Both of these improve in data volume as well, with little to no performance penalties. Reductions in data volume and worst-case edge placement error are on the order of 20-30% and 30-50% respectively compared to a standard baseline design rule set. Luigi Capodieci, Puneet Gupta 0001, Andrew B. Kahng, Dennis Sylvester, Jie Yang 0010 |
DAC | 3 |
| 2004 | Selective gate-length biasing for cost-effective runtime leakage controlabstractWith process scaling, leakage power reduction has become one of the most important design concerns. Multi-threshold techniques have been used to reduce runtime leakage power without sacrificing performance. In this paper, we propose small biases of transistor gate-length to further minimize power in a manufacturable manner. Unlike multi-V th techniques, gate-length biasing requires no additional masks and may be performed at any stage in the design process.Our results show that gate-length biasing effectively reduces leakage power by up to 25% with less than 4% delay penalty. We show the feasibility of the technique in terms of manufacturability and pin-compatibility for post-layout power optimization. We also show up to 54% reduction in leakage uncertainty due to inter-die process variation in circuits when biased gate-lengths, versus only unbiased one, are used. Circuits selectively biased show much less sensitivity to both intra and inter die variations. Puneet Gupta 0001, Andrew B. Kahng, Dennis Sylvester |
DAC | 2 |
| 2004 | Placement feedback: a concept and method for better min-cut placementsabstractThe advent of strong multi-level partitioners has made topdown min-cut placers a favored choice for modern placer implementations. We examine terminal propagation, an important step in min-cut placers, because it is responsible for translating partitioning results into global placement wirelength assumptions. In this work, we identify a previously overlooked problem - ambiguous terminal propagation - and propose a solution based on the concept of feedback from automatic control systems. Implementing our approach in Capo (version 8.7 [5, 10]) and applying it to standard benchmark circuits yields up to 14% wirelength reductions for the IBM benchmarks and 10% reductions for PEKO instances. Experiments also show consistent improvements for routed wirelength, yielding up to 9% wirelength reductions with practical increase in placement runtime. In addition, our method significantly improves routability without building congestion maps, and reduces the number of vias. Andrew B. Kahng, Sherief Reda |
DAC | 1 |
| 2004 | Boosting: Min-Cut Placement with Improved Signal DelayabstractIn this work we improve top-down min-cut placers in the context of timing closure. Using the concept of boosting factors, we adjust net weights according to net spans, so as to reduce the quadratic wirelength. Our method is generic and does not involve any timing analysis during or prior to placement. In essence, we skew the netlength distribution produced by a min-cut placer so as to decrease the number of long nets, with minimal impact on the overall wirelength. Empirically this approach does not significantly affect runtime, but reduces the worst negative slack and total negative slack of industrial benchmarks by up to 70% compared to Capo and a leading industrial placer. Andrew B. Kahng, Igor L. Markov, Sherief Reda |
DATE | 1 |
| 2004 | On legalization of row-based placementsabstractCell overlaps in the results of global placement are guaranteed to prevent successful routing. However, common techniques for fixing these problems may endanger routing in a different way --- through increased wirelength and congestion. We evaluate several such techniques with routability of row-based placements in mind, and propose new ones that, in conjunction with our detail placer, improve overall routability and routed wirelength. Our generic two-phase approach for resolving illegal placements calls for (i) balancing the numbers of cells in rows, (ii) removing overlaps within rows through a generic dynamic programming procedure. Relevant objectives include minimum total perturbation, minimum wirelength increase and minimum maximum movement. Additionally, we trace cell overlaps in min-cut placement to vertical cuts and show that, if bisection cut directions are varied, overlaps anti-correlate with improved wirelength.Empirical validation is performed using placers Capo and Cadence QPlace, followed by various legalizers and detail placers, with subsequent routing by Cadence WarpRoute. We use a number of IBMv2 benchmarks with routing information. Our legalizer reduces both Capo and QPlace placements' wirelength by up to 4% compared to results of Capo legalized by Cadence's QPlace in the ECO mode. Andrew B. Kahng, Igor L. Markov, Sherief Reda |
ACM Great Lakes Symposium on VLSI | 1 |
| 2004 | An analytic placer for mixed-size placement and timing-driven placementabstractWe extend the APlace wirelength-driven standard-cell analytic placement framework of A.A. Kennings and I.L. Markov (2002) to address timing-driven and mixed-size ("boulders and dust") placement. Compared with timing-driven industry tools, evaluated by commercial detailed routing and STA, we achieve an average of 8.4% reduction in cycle time and 7.5% reduction in wirelength for a set of six industry testcases. For mixed-size placement, we achieve an average of 4% wirelength reduction on ISPD02 mixed-size placement benchmarks compared to results of the leading-edge solver, Feng Shui (v2.4) (Khatkhate et al., 2004). We are currently evaluating our placer on industry testcases that combine the challenges of timing constraints, large instance sizes, and embedded blocks (both fixed and unfixed). Andrew B. Kahng, Qinke Wang |
ICCAD | 1 |
| 2004 | Multi-project reticle floorplanning and wafer dicingabstractMulti-project Wafers (MPW) are an efficient way to share the rising costs of mask tooling between multiple prototype and low production volume designs. Packing the different die images on a multi-project reticle leads to new and highly challenging floorplanning formulations, characterized by unusual constraints and complex objective functions. In this paper we study multi-project reticle floorplanning and wafer dicing problems under the prevalent side-to-side wafer dicing technology. Our contributions include practical mathematical programming algorithms and efficient heuristics based on interval-graph coloring which find side-to-side wafer dicing plans with maximum yield for a fixed multi-project reticle floorplan and given per-die maximum dicing margins. We also give novel shelf packing and simulated annealing reticle floorplanning algorithms for maximizing wafer-dicing yield. Experimental results show that our algorithms improve wafer-dicing yield significantly compared to existing industry tools and academic min-area floorplanners. Andrew B. Kahng, Ion I. Mandoiu, Qinke Wang, Xu Xu 0001, Alex Zelikovsky |
ISPD | 1 |
| 2004 | Implementation and extensibility of an analytic placerabstractAutomated cell placement is a critical problem in VLSI physical design. New analytical placement methods that simultaneously spread cells and optimize wirelength have recently received much attention from both academia and industry. A novel and simple objective function for spreading cells over the placement area is described in the patent of Naylor et al. [26]. When combined with a wirelength objective function, this allows efficient simultaneous cell spreading and wirelength optimization using nonlinear optimization techniques. In this work, we implement an analytic placer (APlace) according to these ideas (which have other precedents in the open literature), and conduct in-depth analysis of characteristics and extensibility of the placer. Our contributions are as follows. (1) We perform analysis and empirical studies of relevant characteristics of the objective functions described in [26]. (2) We extend the objective functions with congestion information. (3) We implement a top-down hierarchical (multilevel) placer (APlace) based on the objective functions. The half-perimeter wirelength of APlace outperforms that of Cadence QPlace (SE5.4), UCLA Dragon (v3.01) and Capo (v8.7) respectively by 6.8%, 2.6% and 6.5% on average. When these placements are detail-routed using Cadence WRoute (SE5.4), the average improvement in final wirelength is 8.2%, 4.2% and 10.4% over QPlace, Dragon and Capo, respectively. (4) We extend the placer to perform I/O-core co-placement. I/Os can be evenly distributed without damaging the wirelength figure of merit. (5) We also extend the placer to handle constraints for mixed-signal designs (symmetry, alignment, etc.) and evaluate the impact of such constraints on runtime and wirelength. Andrew B. Kahng, Qinke Wang |
ISPD | 1 |
| 2004 | Effective iterative techniques for fingerprinting design IPabstractFingerprinting is an approach that assigns a unique and invisible ID to each sold instance of the intellectual property (IP). One of the key advantages fingerprinting-based intellectual property protection (IPP) has over watermarking-based IPP is the enabling of tracing stolen hardware or software. Fingerprinting schemes have been widely and effectively used to achieve this goal; however, their application domain has been restricted only to static artifacts, such as image and audio, where distinct copies can be obtained easily. In this paper, we propose the first generic fingerprinting technique that can be applied to an arbitrary synthesis (optimization or decision) or compilation problem and, therefore to hardware and software IPs. The key problem with design IP fingerprinting is that there is a need to generate a large number of structurally unique but functionally and timing identical designs. To reduce the cost of generating such distinct copies, we apply iterative optimization in an incremental fashion to solve a fingerprinted instance. Therefore, we leverage on the optimization effort already spent in obtaining previous solutions, yet we generate a uniquely fingerprinted new solution. This generic approach is the basis for developing specific fingerprinting techniques for four important problems in VLSI CAD: partitioning, graph coloring, satisfiability, and standard-cell placement. We demonstrate the effectiveness of the new fingerprinting-based IPP techniques on a number of standard benchmarks. Andrew E. Caldwell, Hyun-Jin Choi, Andrew B. Kahng, Stefanus Mantik, Miodrag Potkonjak, Gang Qu 0001, Jennifer Wong-Ma |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2004 | Nontree routing for reliability and yield improvement [IC layout]abstractWe propose to introduce redundant interconnects for manufacturing yield and reliability improvement. By introducing redundant interconnects, the potential for open faults is reduced, at the cost of increased potential for short faults. Overall, manufacturing yield and fault tolerance can be improved. We focus on a postprocessing, tree-augmentation approach, which can be easily integrated in current physical design flows. Our contributions are as follows. 1) We formulate the problem as a variant of the classical two-edge-connectivity augmentation problem in which we take into account such practical issues as wirelength increase budget, routing obstacles, and the use of Steiner points. 2) We show that an optimum solution can always be found on the Hanan grid defined by the terminals and the corners of the feasible routing region. 3) We give a compact integer program formulation which is solved in practical runtime by the commercial optimization package CPLEX for nets with up to 100 terminals. 4) We give a well-scaling greedy algorithm which has a practical runtime up to 1000 terminals, and comes on the average within 1%-2% of the optimum computed by CPLEX. 5). We give a comprehensive experimental study comparing the solution quality and runtime of our methods with the best methods reported in the literature. Andrew B. Kahng, Bao Liu 0001, Ion I. Mandoiu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2004 | Local unidirectional bias for cutsize-delay tradeoff in performance-driven bipartitioningabstractTraditional multilevel partitioning approaches have shown good performance with respect to cutsize, but offer no guarantees with respect to system performance. Timing-driven partitioning methods based on iterated net reweighting, partitioning, and timing analysis have been proposed (Ababei et al., 2002), as well as methods that apply degrees of freedom such as retiming (Cong et al., 2000), (Cong et al., 2002). In this paper, we identify and validate a simple approach to timing-driven partitioning based on the concept of "V-shaped nodes." We observe that the presence of V-shaped nodes can badly impact circuit performance, as measured by maximum hopcount across the cutline or similar path delay criteria. We extend traditional the Fiduccia-Mattheyses (FM) variant of the Kernighan-Lin (Kernighan and Lin, 1970) algorithm approaches to directly eliminate or minimize "distance-k V-shaped nodes" in the bipartitioning solution, achieving an attractive tradeoff between cutsize and path delay. Experiments show that in comparison to MLPart (Caldwell et al., 2000), our method can reduce the maximum hopcount by 39% while only slightly increasing cutsize and runtime. No previous method improves path delay in such a transparent manner. The new partitioner is incorporated into a placer (http://vlsicad.ucsd.edu/GSRC/bookshelf/Slots/Placement/Capo/) and circuit delay is evaluated by a commercial static timing analyzer (http://www.ece.uci.edu/eceware/cadence/spl I.bar/docs/pearluser/). The empirical results show that the delay is significantly reduced, at the cost of very acceptable impacts on wirelength and runtime. Andrew B. Kahng, Xu Xu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2003 | Routing-aware scan chain orderingabstractScan chain insertion can have a large impact on routability, wirelength and timing of the design. We present a routing-driven methodology for scan chain ordering with minimum wirelength objective. A routing-based approach to scan chain ordering, while potentially more accurate, can result in TSP (Traveling Salesman Problem) instances which are asymmetric and highly non-metric; this may require a careful choice of solvers. We evaluate our new methodology on recent industry place-and-route blocks with 1200 to 5000 scan cells. We show substantial wirelength reductions for the routing-based flow, versus the traditional placement-based flow: in a number of our test cases, over 86% of scan routing overhead is saved. Even though our experiments are so far timing-oblivious, the routing-based flow does also improve evaluated timing, and practical timing-driven extensions appear feasible. Puneet Gupta 0001, Andrew B. Kahng, Stefanus Mantik |
ASP-DAC | 2 |
| 2003 | Highly scalable algorithms for rectilinear and octilinear Steiner treesabstractThe rectilinear Steiner minimum tree (RSMT) problem, which asks for a minimum-length interconnection of a given set of terminals in the rectilinear plane, is one of the fundamental problems in electronic design automation. Recently there has been renewed interest in this problem due to the need for highly scalable algorithms able to handle nets with tens of thousands of terminals. In this paper we give a practical O (n log2 n) heuristic for computing near-optimal rectilinear Steiner trees based on a batched version of the greedy triple contraction algorithm of Zelikovsky [21]. Experiments conducted on both random and industry testcases show that our heuristic matches or exceeds the quality of best known RSMT heuristics, e.g., on random instances with more than 100 terminals our heuristic improves over the rectilinear minimum spanning tree by an average of 11%. Moreover, our heuristic has very well scaling runtime, e.g., it can route a 34k-terminals net extracted from a real design in less than 25 seconds compared to over 86 minutes needed by the O(n2) edge-based heuristic of Borah, Owens, and Irwin [3]. Since our heuristic is graph-based, it can be easily modified to handle practical considerations such as routing obstacles, preferred directions, via costs, and octilinear routing - indeed, experimental results show only a small factor increase in runtime when switching from rectilinear to octilinear routing. Andrew B. Kahng, Ion I. Mandoiu, Alex Zelikovsky |
ASP-DAC | 1 |
| 2003 | An algebraic multigrid solver for analytical placement with layout based clusteringabstractAn efficient matrix solver is critical to the analytical placement. As the size of the matrix becomes huge, the multilevel methods turn out to be more efficient and more scalable. Algebraic Multigrid (AMG) is a multilevel technique to speedup the iterative matrix solver [10]. We apply the algebraic multigrid method to solve the linear equations that arise from the analytical placement. A layout based clustering scheme is put forward to generate coarsening levels for the multigrid method. The experimental results show that the algebraic multigrid solver is promising for analytical placement. Hongyu Chen 0001, Chung-Kuan Cheng, Nan-Chi Chou, Andrew B. Kahng, John F. MacDonald, Peter Suaris, Bo Yao 0004, Zhengyong Zhu |
DAC | 4 |
| 2003 | Performance-impact limited area fill synthesisabstractChemical-mechanical planarization (CMP) and other manufacturing steps in very deep-submicron VLSI have varying effects on device and interconnect features, depending on the local layout density. To improve manufacturability and performance predictability, area fill features are inserted into the layout to improve uniformity with respect to density criteria. However, the performance impact of area fill insertion is not considered by any fill method in the literature. In this paper, we first review and develop estimates for capacitance and timing overhead of area fill insertions. We then give the first formulations of the Performance Impact Limited Fill (PIL-Fill) problem with the objective of either minimizing total delay impact (MDFC) or maximizing the minimum slack of all nets (MSFC), subject to inserting a given prescribed amount of fill. For the MDFC PIL-Fill problem, we describe three practical solution approaches based on Integer Linear Programming (ILP-I and ILP-II) and the Greedy method. For the MSFC PIL-Fill problem, we describe an iterated greedy method that integrates call to an industry static timing analysis tool. We test our methods on layout testcases obtained from industry. Compared with the normal fill method [3], our ILP-II method for MDFC PIL-Fill problem achieves between 25-% and 90% reduction in terms of total weighted edge delay (roughly, a measure of sum of node slacks) impact while maintaining identical quality of the layout density control; and our iterated greedy method for MSFC PIL-Fill problem also shows significant advantage with respect to the minimum slack of nets on post-fill layout. Yu Chen 0005, Puneet Gupta 0001, Andrew B. Kahng |
DAC | 3 |
| 2003 | A cost-driven lithographic correction methodology based on off-the-shelf sizing toolsabstractAs minimum feature sizes continue to shrink, patterned features have become significantly smaller than the wavelength of light used in optical lithography. As a result, the requirement for dimensional variation control, especially in critical dimension (CD) 3σ, has become more stringent. To meet these requirements, resolution enhancement techniques (RET) such as optical proximity correction (OPC) and phase shift mask (PSM) technology are applied. These approaches result in a substantial increase in mask costs and make the cost of ownership (COO) a key parameter in the comparison of lithography technologies. No concept of function is injected into the mask flow; that is, current OPC techniques are oblivious to the design intent, and the entire layout is corrected uniformly with the same effort. We propose a novel minimum cost of correction (MinCorr) methodology to determine the level of correction for each layout feature such that prescribed parametric yield is attained with minimum total RET cost. We highlight potential solutions to the MinCorr problem and give a simple mapping to traditional performance optimization. We conclude with experimental results showing that substantial RET costs may be saved while maintaining a given desired level of parametric yield. Puneet Gupta 0001, Andrew B. Kahng, Dennis Sylvester, Jie Yang 0010 |
DAC | 2 |
| 2003 | Nanometer design: place your betsabstractOverview Two years ago, DAC-2001 attendees enjoyed a thrilling debatepanel, “Who’s Got Nanometer Design Under Control?”, pitting sky-is-falling Physics die-hards against not-to-worry Methodology gurus. Then, the DAC audience overwhelmingly voted the match for the Methodologists. Now, we've just gone through the biggest business downturn in the industry's history, and we're hearing more and more about chip failures due to 130nm physical effects. Both physics and economics are a lot worse than we thought two years ago. Where are those simple, correct-by-construction methodologies for signal integrity, power integrity, low-power, etc. that we were promised? Were we bamboozled by glib promises from those Methodologists? In this session, we bring back the panelists from two years ago, not for another debate, but to hear well-reasoned perspectives on how to prioritize spending to address nanometer design challenges. Yes, methodology can solve any problem – but now we want to know which problems, in what priority order, at what cost. The panel will address the following questions. • What are the economic impacts and significance of the key nanometer design challenges, relative to each other? • Which nanometer design problems merit responsible R&D investment, in what amounts and proportion? • What is the likelihood of success, both near-term and longterm, in solving key nanometer design challenges? • Where will the answers come from? To keep the discussion very concrete, each panelist will be given a $100 budget, and must defend their allocation of this budget to attack various design problems. Where should the $100 be spent? The audience will determine the best-reasoned allocation, and the winning panelist keeps all the money. Andrew B. Kahng, Shekhar Borkar, John M. Cohn, Antun Domic, Patrick Groeneveld, Louis K. Scheffer, Jean-Pierre Schoellkopf |
DAC | 1 |
| 2003 | Area Fill Generation With Inherent Data Volume Reduction
Yu Chen 0005, Andrew B. Kahng, Gabriel Robins, Alex Zelikovsky |
DATE | 2 |
| 2003 | A Novel Metric for Interconnect Architecture Performance
Parthasarathi Dasgupta, Andrew B. Kahng, Swamy Muddu |
DATE | 2 |
| 2003 | The Y-Architecture for On-Chip Interconnect: Analysis and Methodology
Hongyu Chen 0001, Chung-Kuan Cheng, Andrew B. Kahng, Ion I. Mandoiu, Qinke Wang, Bo Yao 0004 |
ICCAD | 3 |
| 2003 | Manufacturing-Aware Physical Design
Puneet Gupta 0001, Andrew B. Kahng |
ICCAD | 2 |
| 2003 | Layout-Aware Scan Chain Synthesis for Improved Path Delay Fault Coverage
Puneet Gupta 0001, Andrew B. Kahng, Ion I. Mandoiu |
ICCAD | 2 |
| 2003 | Evaluation of Placement Techniques for DNA Probe Array Layout
Andrew B. Kahng, Ion I. Mandoiu, Sherief Reda, Xu Xu 0001, Alex Zelikovsky |
ICCAD | 1 |
| 2003 | Design Flow Enhancements for DNA ArraysabstractDNA probe arrays have recently emerged as one of the core genomic technologies. Exploiting analogies between manufacturing processes for DNA arrays and for VLSI chips, we demonstrate the potential for transfer of methodologies from the 40-year old field of electronic design automation to the newer DNA array design field. Our main contributions are the following. (1) We give a new design flow for DNA arrays which enhances current methodologies by adding flow-awareness to each optimization step and introducing feedback loops. (2) We propose solution methods for new formulations integrating multiple design steps, including probe selection, placement, and embedding. (3) We give results of a comprehensive experimental study showing that significant improvements in solution quality can be achieved by using the enhanced methodologies. Andrew B. Kahng, Ion I. Mandoiu, Sherief Reda, Xu Xu 0001, Alex Zelikovsky |
ICCD | 1 |
| 2003 | Research directions for coevolution of rules and routersabstractDesign rules in advanced IC manufacturing processes are increasingly problematic for modern router architectures and algorithms. This paper first reviews types and causes of "difficult" design rules, as well as implications for current routing approaches. Next, some basic router components are assessed with respect to future viability. Last, the paper discusses prospects for future "coevolution" of design rules and detailed routing methods. Andrew B. Kahng |
ISPD | 1 |
| 2003 | Local unidirectional bias for smooth cutsize-delay tradeoff in performance-driven bipartitioningabstractTraditional multilevel partitioning approaches have shown good performance with respect to cutsize, but offer no guarantees with respect to system performance. Timing-driven partitioning methods based on iterated net reweighting, partitioning and timing analysis have been proposed [2], as well as methods that apply degrees of freedom such as retiming [7] [5]. In this work, we identify and validate a simple approach to timing-driven partitioning, based on the concept of “V-shaped nodes”. We observe that the presence of V-shaped nodes can badly impact circuit performance, as measured by maximum hopcount across the cutline or similar path delay criteria. We extend traditional KLFM approaches to directly eliminate or minimize “distance-k Vshaped nodes ” in the bipartitioning solution, achieving a smooth trade-off between cutsize and path delay. Experiments show that in comparison to MLPart [4], our method can reduce the maximum hopcount by 39 % while only slightly increasing cutsize and runtime. No previous method improves path delay in such a transparent manner. Andrew B. Kahng, Xu Xu 0001 |
ISPD | 1 |
| 2003 | Engineering a scalable placement heuristic for DNA probe arraysabstractDesign of DNA arrays for very large-scale immobilized polymer synthesis (VLSIPS) [8] seeks to minimize effects of unintended illumination during mask exposure steps. [9, 14] formulate this requirement as the Border Minimization Problem and give methods for placement (at array sites) and embedding (in the mask sequence) of probes in both synchronous and asynchronous regimes. These previous methods do not address several practical details of the application and, more critically, are not scalable to the O(108) probes contemplated for next-generation probe arrays. In this work, we make two main contributions: Andrew B. Kahng, Ion I. Mandoiu, Pavel A. Pevzner, Sherief Reda, Alex Zelikovsky |
RECOMB | 1 |
| 2003 | On the skew-bounded minimum-buffer routing tree problemabstractBounding the load capacitance at gate outputs is a standard element in today's electrical correctness methodologies for high-speed digital very large scale integration design. Bounds on load caps improve coupling-noise immunity, reduce degradation of signal transition edges, and reduce delay uncertainty due to coupling noise (Kahng et al. 1998). For clock and test distribution, an additional design requirement is bounding the buffer skew, i.e., the difference between the maximum and the minimum number of buffers over all of the source-to-sink paths in the routing tree, since buffer skew is one of the main factors affecting delay skew (Tellez and Sarrafzadeh 1997). In this paper, we consider algorithms for buffering a given tree with the minimum number of buffers under given load cap and buffer skew constraints. We show that the greedy algorithm proposed by Tellez and Sarrafzadeh is suboptimal for nonzero buffer-skew bounds and give examples showing that no bottom-up greedy algorithm can achieve optimality. The main contribution of the paper is an optimal dynamic programming algorithm for the problem. Experiments on test cases extracted from recent industrial designs show that the dynamic programming algorithm has practical running time and saves up to 37.5% of the buffers inserted by Tellez and Sarrafzadeh's algorithm. Christoph Albrecht, Andrew B. Kahng, Bao Liu 0001, Ion I. Mandoiu, Alex Zelikovsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2003 | Minimum buffered routing with bounded capacitive load for slew rate and reliability controlabstractIn high-speed digital VLSI design, bounding the load capacitance at gate outputs is a well-known methodology to improve coupling noise immunity, reduce degradation of signal transition edges, and reduce delay uncertainty due to coupling noise. Bounding load capacitance also improves reliability with respect to hot-carrier oxide breakdown and AC self-heating in interconnects, and guarantees bounded input rise/fall times at buffers and sinks. This paper introduces a new minimum-buffer routing problem (MBRP) formulation which requires that the capacitive load of each buffer, and of the source driver, be upper-bounded by a given constant. Our contributions are as follows: We give linear-time algorithms for optimal buffering of a given routing tree with a single (inverting or noninverting) buffer type. For simultaneous routing and buffering with a single noninverting buffer type, we prove that no algorithm can guarantee a factor smaller than 2 unless P=NP and give an algorithm with approximation factor slightly larger than 2 for typical buffers. For the case of a single inverting buffer type, we give an algorithm with approximation factor slightly larger than 4. We give local-improvement and clustering based MBRP heuristics with improved practical performance, and present a comprehensive experimental study comparing the runtime/quality tradeoffs of the proposed MBRP heuristics on test cases extracted from recent industrial designs. Charles J. Alpert, Andrew B. Kahng, Bao Liu 0001, Ion I. Mandoiu, Alex Zelikovsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2003 | Hierarchical whitespace allocation in top-down placementabstractIncreased transistor density in modern commercial ICs typically originates in new manufacturing and defect prevention technologies. Additionally, better utilization of such low-level transistor density may result from improved software that makes fewer assumptions about physical layout in order to reliably automate the design process. In particular, recent layouts tend to have large amounts of whitespace, which is not handled properly by older tools. We observe that a major computational difficulty arises in partitioning-driven top-down placement when regions of a chip lack whitespace. This tightens balance constraints for min-cut partitioning and hampers move-based local-search heuristics such as Fiduccia-Mattheyses. However, the local lack of whitespace is often caused by very unbalanced distribution of whitespace during previous partitioning, and this concern is emphasized in chips with large overall whitespace. This paper focuses on accurate computation of tolerances to ensure smooth operation of common move-based iterative partitioners, while avoiding cell overlaps. We propose a mathematical model of hierarchical whitespace allocation in placement, which results in a simple computation of partitioning tolerance purely from relative whitespace in the block and the number of rows in the block. Partitioning tolerance slowly increases as the placer descends to lower levels, and relative whitespace in all blocks is limited from below (unless partitioners return "illegal" solutions), thus preventing cell overlaps. This facilitates good use of whitespace when it is scarce and prevents very dense regions when large amounts of whitespace are available. Our approach improves the use of the available whitespace during global placement, thus leading to smaller whitespace requirements. Existing techniques, particularly those based on simulated annealing, can be applied after global placement to bias whitespace with respect to particular concerns, such as routing congestion, heat dissipation, crosstalk noise and DSM yield improvement. Andrew E. Caldwell, Andrew B. Kahng, Igor L. Markov |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2003 | Improved a priori interconnect predictions and technology extrapolation in the GTX systemabstractA priori interconnect prediction and technology extrapolation are closely intertwined. Interconnect predictions are at the core of technology extrapolation models of achievable system power, area density, and speed. Technology extrapolation, in turn, informs a priori interconnect prediction via models of interconnect technology and interconnect optimizations. In this paper, we address the linkage between a priori interconnect prediction and technology extrapolation in two ways. First, we describe how rapid changes in technology, as well as rapid evolution of prediction methods, require a dynamic and flexible framework for technology extrapolation. We then develop a new tool, the GSRC technology extrapolation system (GTX), which allows capture of such knowledge and rapid development of new studies. Second, we identify several "nontraditional" facets of interconnect prediction and quantify their impact on key technology extrapolations. In particular, we explore the effects of interconnect design optimizations such as shield insertion, repeater sizing and repeater staggering, as well as modeling choices for RLC interconnects. Yu Cao 0001, Chenming Hu, Xuejue Huang, Andrew B. Kahng, Igor L. Markov, Michael Oliver, Dirk Stroobandt, Dennis Sylvester |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2002 | Tools or users: which is the bigger bottleneck?abstractAs chip design becomes ever more complex, fewer design teams are succeeding. Who's to blame? On one hand, tools are hard to use, buggy, not interoperable, and have missing functionality. On the other hand, there is a wide range of engineering skills within the user population, and tools can be abused within flawed methodologies. This panel will quantify and prioritize the key gaps, including interoperability, that must be addressed on both sides. Andrew B. Kahng, Ronald Collett, Patrick Groeneveld, Lavi Lev, Nancy Nettleton, Paul K. Rodman, Lambert van den Hoven |
DAC | 1 |
| 2002 | Non-tree routing for reliability and yield improvementabstractWe propose to introduce redundant interconnects for manufacturing yield and reliability improvement. By introducing redundant interconnects, the potential for open faults is reduced at the cost of increased potential for short faults; overall, manufacturing yield and fault tolerance can be improved. We focus on a post-processing, tree augmentation approach which can be easily integrated in current physical design flows. Our contributions are as follows:• We formulate the problem as a variant of the classical 2-edge-connectivity augmentation problem in which we take into account such practical issues as wirelength increase budget, routing obstacles, and use of Steiner points.• We show that an optimum solution can always be found on the Hanan grid defined by the terminals and the corners of the feasible routing region.• We give a compact integer program formulation which, for up to 100 terminal nets, is solved in practical runtime by the commercial optimization package CPLEX.• We give a well-scaling greedy algorithm which has practical runtime up to 1,000 terminals, and comes on the average within 1--2% of the optimum computed by CPLEX.• We give a comprehensive experimental study comparing the solution quality and runtime of our methods with the best reported methods for 2-edge-connectivity augmentation, including a sophisticated heuristic based on minimum-weight branchings [9] and a recent genetic algorithm [14].Experiments on randomly generated and industry testcases show that our greedy augmentation method achieves significant increase in reliability (as measured by the percentage of biconnected tree edges) with very small increase in wirelength. For example, on 1,000 terminal nets the average percentage of biconnected tree edges is 34.19% for a wire-length increase of only 1%, and 87.73% for a wirelength increase of 20%. SPICE simulations on industry routed nets show that non-tree routing has the additional benefit of reducing maximum sink delay by an average of 28.26% compared to Steiner routing, and by an average of 3.72% compared to timing optimized routing. SPICE simulations further imply that non-tree routing has smaller delay variation due to process variability. Andrew B. Kahng, Bao Liu 0001, Ion I. Mandoiu |
ICCAD | 1 |
| 2002 | Closing the smoothness and uniformity gap in area fill synthesisabstractControl of variability in the back end of the line, and hence in interconnect performance as well, has become extremely difficult with the introduction of new materials such as copper and low-k dielectrics. Uniformity of chemical-mechanical planarization (CMP) requires the addition of area fill geometries into the layout, in order to smoothen the variation of feature densities across the die. Our work addresses the following smoothness gap in the recent literature on area fill synthesis. (1)The very first paper on the filling problem (Kahng et al., ISPD98 [7]) noted that there is potentially a large difference between the optimum window densities in fixed dissections vs. when all possible windows in the layout are considered. (2)Despite this observation, all filling methods since 1998 minimize and evaluate density variation only with respect to a fixed dissection. This paper gives the first evaluation of existing filling algorithms with respect to "gridless" ("floating-window") mode, according to both the effective and spatial density models. Our experiments indicate surprising advantages of Monte-Carlo and greedy strategies over "optimal" linear programming (LP) based methods. Second, we suggest new, more relevant methods of measuring a local uniformity based on Lipschitz conditions, and empirically demonstrate that Monte-Carlo methods are inherently better than LP with respect to the new criteria. Finally, we propose new LP-based filling methods that are directly driven by the new criteria, and show that these methods indeed help close the "smoothness gap". Yu Chen 0005, Andrew B. Kahng, Gabriel Robins, Alex Zelikovsky |
ISPD | 2 |
| 2002 | A roadmap and vision for physical designabstractThis invited paper offers "roadmap and vision" for physical design. The main messages are as follows. (1) The high-level roadmap for physical design is static and well-known. (2) Basic problems remain untouched by fundamental research. (3) Academia should not overemphasize back- filling and formulation over innovation and optimization. (4) The physical design field must become more mature and efficient in how it prioritizes research directions and uses its human resources. (5) The scope of physical design must expand (up to package and system, down to manufacturing interfaces, out to novel implementation technologies, etc.), even as renewed focus is placed on basic optimization technology. Andrew B. Kahng |
ISPD | 1 |
| 2002 | Min-max placement for large-scale timing optimizationabstractAt the 250nm technology node, interconnect delays account for over 40% of worst delays [12]. Transition to 130nm and below increases this figure, and hence the relative importance of timing-driven placement for VLSI. Our work introduces a novel minimization of maximal path delay that improves upon previously known algorithms for timing-driven placement. Our placement algorithms have provable properties and are fast in practice. Empirical validation is based on extending a scalable min-cut placer with proven quality in wirelength- and congestion-driven placement [4]. The CPU overhead of the timing-driven capability is within 50%. We placed industrial circuits and evaluated the resulting layouts with a commercial static timing analyzer. Andrew B. Kahng, Stefanus Mantik, Igor L. Markov |
ISPD | 1 |
| 2002 | Border Length Minimization in DNA Array Design
Andrew B. Kahng, Ion I. Mandoiu, Pavel A. Pevzner, Sherief Reda, Alex Zelikovsky |
WABI | 1 |
| 2002 | Area fill synthesis for uniform layout densityabstractChemical-mechanical polishing (CMP) and other manufacturing steps in very deep submicron very large scale integration have varying effects on device and interconnect features, depending on local characteristics of the layout. To improve manufacturability and performance predictability, the authors seek to make a layout uniform with respect to prescribed density criteria, by inserting "area fill" geometries into the layout. In this paper, they make the following contributions. First, the authors define the flat, hierarchical, and multiple-layer filling problems, along with a unified density model description. Secondly, for the flat filling problem, they summarize current linear programming approaches with two different objectives, i.e., the Min-Var and Min-Fill objectives. They then propose several new Monte Carlo-based filling methods with fast dynamic data structures. Thirdly, they give practical iterated methods for layout density control for CMP uniformity based on linear programming, Monte Carlo, and greedy algorithms. Fourthly, to address the large data volume and inherent lack of scalability of flat layout density control, the authors propose practical methods for hierarchical layout density control. These methods smoothly trade off runtime, solution quality, and output data volume. Finally, they extend the linear programming approaches and present new Monte Carlo-based methods for the multiple-layer filling problem. Comparisons with previous filling methods show the advantages of the new iterated Monte Carlo and iterated greedy methods for both flat and hierarchical layouts and for both density models (spatial density and effective density). The authors achieve near-optimal filling for flat layouts with respect to each of these objectives. Their experiments indicate that the hybrid hierarchical filling approach is efficient, scalable, accurate, and highly competitive with existing methods (e.g., linear programming-based techniques) for hierarchical layouts. Yu Chen 0005, Andrew B. Kahng, Gabriel Robins, Alex Zelikovsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2002 | Provably good global buffering by generalized multiterminalmulticommodity flow approximationabstractTo implement high-performance global interconnect without impacting the placement and performance of existing blocks, the use of buffer blocks is becoming increasingly popular in structured-custom and block-based application specified integrated circuit methodologies. We address the problem of how to perform the buffering of global multiterminal nets given an existing buffer block plan. We give provably good and heuristic algorithms for this problem. The method routes connections using available buffer blocks, such that required upper and lower bounds on buffer intervals are satisfied. In addition, the algorithms allow more than one buffer to be inserted into any given connection and observe upper bounds and parity constraints on the number of buffers per connection. Most importantly, and unlike previous works on the problem, we take into account: 1) multiterminal nets; 2) multiple routing layers; 3) simultaneous buffered routing and compaction; and 4) buffer libraries. Our method outperforms existing algorithms for the problem, based on two-pin decompositions of the nets, and has been validated on top-level layouts extracted from a recent high-end microprocessor design. Feodor F. Dragan, Andrew B. Kahng, Ion I. Mandoiu, Sudhakar Muddu, Alex Zelikovsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2002 | Toward better wireload models in the presence of obstaclesabstractWirelength estimation techniques typically contain a site density function that enumerates all possible path sites for each wirelength in an architecture and an occupation probability function that assigns a probability to each of these paths to be occupied by a wire. In this paper, we apply a generating polynomial technique to derive complete expressions for site density functions which take effects of layout region aspect ratio and the presence of obstacles into account. The effect of an obstacle is separated into two parts: the terminal redistribution effect and the blockage effect. The layout region aspect ratio and the obstacle area are observed to have a much larger effect on the wirelength distribution than the obstacle's aspect ratio and location. Accordingly, we suggest that these two parameters be included as indices of lookup tables in wireload models. Our results apply to a priori wirelength estimation schemes in chip planning tools to improve parasitic estimation accuracy and timing closure; this is particularly relevant for system-on-chip designs where IP blocks are combined with row-based layout. Chung-Kuan Cheng, Andrew B. Kahng, Bao Liu 0001, Dirk Stroobandt |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2001 | Hierarchical dummy fill for process uniformityabstractTo improve manufacturability and performance predictability, we seek to make a layout uniform with respect to prescribed density criteria, by inserting "fill" geometries into the layout. Previous approaches for at layout density control are not scalable due to the necessity of solving very large linear programs, the large data volume of the solution, and the impact of hierarchy-breaking on verification. In this paper, we give the first methods for hierarchical layout density control for process uniformity. Our approach trades off naturally between runtime, solution quality, and output data volume. We also allow generation of compressed GDSII of fill geometries. Our experiments show that this hybrid hierarchical filling approach saves data volume and is scalable, while yielding solution quality that is competitive with existing Monte-Carlo and linear programming based approaches. Yu Chen 0005, Andrew B. Kahng, Gabriel Robins, Alex Zelikovsky |
ASP-DAC | 2 |
| 2001 | Toward better wireload models in the presence of obstaclesabstractEfficient and accurate interconnect estimation is crucial to design convergence. With System-on-Chip design, IP blocks form routing obstacles that cannot be accounted for by existing a priori wirelength estimations. In this paper, we identify two distinct effects of obstacles on interconnection length: (i) changes due to the redistribution of interconnect terminals and (ii) detours that have to be made around the obstacles. Theoretical expressions of both effects for point-to-point nets with a single obstacle are derived and compared to experimental observations. We also experimentally assess these effects for multi-terminal interconnections and in the presence of multiple obstacles. We single out cases where the effects are additive, which suggests the use of lookup tables and equivalent blockage relations. Our results are applicable in chip planning tools, where they enable improved accounting for obstacles in a priori wirelength estimation schemes. Chung-Kuan Cheng, Andrew B. Kahng, Bao Liu 0001, Dirk Stroobandt |
ASP-DAC | 2 |
| 2001 | Provably good global buffering by multi-terminal multicommodity flow approximationabstractTo implement high-performance global interconnect without impacting the placement and performance of existing blocks, the use of buffer blocks is becoming increasingly popular in structured-custom and block-based ASIC methodologies. Recent works by Cong, Kong and Pan [5] and Tang and Wong [18] give algorithms to solve the buffer block planning problem. In this paper, we address the problem of how to perform buffering of global multiterminal nets given an existing buffer block plan. We give a provably good algorithm based on a recent approach of Garg and Könemann [8] and Fleischer [7] (see also Albrecht [1] and Dragan et al. [6]). Our method routes connections using available buffer blocks, such that required upper and lower bounds on buffer intervals - as well as wirelength upper bounds per connection - are satisfied. In addition, our algorithm allows more than one buffer to be inserted into any given connection and observes buffer parity constraints. Most importantly, and unlike previous works on the problem [5, 18, 6], we take into account multiterminal nets. Our algorithm outperforms existing algorithms for the problem [5, 6], which are based on 2-pin decompositions of the nets. The algorithm has been validated on top-level layouts extracted from a recent high-end microprocessor design. Feodor F. Dragan, Andrew B. Kahng, Ion I. Mandoiu, Sudhakar Muddu, Alex Zelikovsky |
ASP-DAC | 2 |
| 2001 | Design technology productivity in the DSM era (invited talk)abstractFuture requirements for design technology are always uncertain due to changes in process technology, system implementation platforms, and applications markets. To correctly identify the design technology need, and to deliver this technology at the right time, the design technology community - commercial vendors, captive CAD organizations, and academic researchers - must focus on improving design technology time-to-market and quality-of-result. Put another way, we must address the well-known Design Productivity Gap by addressing the Design Technology Productivity Gap. The future of design technology is most uncertain in the physical implementation space, where many complexities must be managed without sacrificing system value or turnaround time. This talk will begin with examples of changes in methodology and design tool infrastructure needed to support future physical implementation. Then, recent initiatives in the MARCO Gigascale Silicon Research Center (GSRC) are described which support the design technology and designer communities in creating the right design technology when it is needed. Andrew B. Kahng |
ASP-DAC | 1 |
| 2001 | New graph bipartizations for double-exposure, bright field alternating phase-shift mask layoutabstractWe describe new graph bipartization algorithms for lay-out modification and phase assignment of bright-field alternating phase-shifting masks (AltPSM) [25]. The problem of layout modification for phase-assignability reduces to the problem of making a certain layout-derived graph bipartite (i.e., 2-colorable). Previous work [3] solves bipartization optimally for the dark field alternating PSMregime. Only one degree of freedom is allowed (and relevant) for such a bipartization: edge deletion, which corresponds to increasing the spacing between features in order to remove phase conflict. Unfortunately, dark-field PSM is used only for contact layers, due to limitations of negative photoresists. Poly and metal layers are actually created using positive photoresists and bright-field masks. In this paper, we define a new graph bipartization formulation that pertains to the more technologically relevant bright-field regime. Previous work [3] does not apply to this regime. This formulation allows two degrees of freedom for layout perturbation: (i) increasing the spacing between features, and (ii) increasing the width of critical features. Each of these corresponds to node deletion in a new layout-derived graph that we define, called the feature graph. Graph bipartization by node deletion asks for a minimum weight node set A such that deletion of A makes the graph bipartite. Unlike bipartization by edge deletion, this problem is NP-hard. We investigate several practical heuristics for the node deletion bipartization of planar graphs, including one that has 9/4 approximation ratio. Computational experience with industrial VLSI layout benchmarks shows promising results. Andrew B. Kahng, Shailesh Vaya, Alex Zelikovsky |
ASP-DAC | 1 |
| 2001 | Beyond the red brick wall (panel): challenges and solutions in 50nm physical designabstractAggressive technology scaling will push us into a 50nm regime within a decade. Most entries in current ITRS for the technology node are painted out in red, indicating "No know solutions". In the physical implementation domain, we are facing severe challenges in various aspects such as interconnect performance degradation, signal integrity, reliability, manufacturing variability, etc. These challenges will continue to grow for the future. In this panel, our panelists will present their own view of the most difficult challenges in the 50nm regime, and possible solutions to break through the red brick wall as well, followed by a live discussion on the approaches we should take for successful 50nm physical implementation. Hidetoshi Onodera, Andrew B. Kahng, Wayne Wei-Ming Dai, Sani R. Nassif, Akira Tanabe, Toshihiro Hattori |
ASP-DAC | 2 |
| 2001 | Panel: Is Nanometer Design Under Control?abstractAs fabrication technology moves to 100 nm and below, profound nanometer effects become critical in developing silicon chips with hundreds of millions of transistors. Both EDA suppliers and system houses have been re- tooling, and new methodologies have been emerging. Will these efforts meet the challenges of nanometer silicon such as performance closure, power, reliability, manufacturability, and cost? Which aspects of nanometer design are, or are not, under control? This session will consist of a debate between two teams of distinguished representatives from EDA suppliers and system design houses. Which side has the right answers and roadmap? You and a panel of judges will decide! Andrew B. Kahng, Bing J. Sheu, Nancy Nettleton, John M. Cohn, Shekhar Borkar, Louis K. Scheffer, Ed Cheng, Sang Wang |
DAC | 1 |
| 2001 | Minimum-Buffered Routing of Non-Critical Nets for Slew Rate and Reliability ControlabstractIn high-speed digital VLSI design, bounding the load capacitance at gate outputs is a well-known methodology to improve coupling noise immunity, reduce degradation of signal transition edges, and reduce delay uncertainty due to coupling noise. Bounding load capacitance also improves reliability with respect to hot-carrier oxide breakdown and AC self-heating in interconnects, and guarantees bounded input rise/fall times at buffers and sinks. This paper introduces a new minimum-buffer routing problem (MBRP) formulation which requires that the capacitive load of each buffer, and of the source driver, be upper-bounded by a given constant. Our contributions include the following. (i) We give linear-time algorithms for optimal buffering of a given routing tree with a single (inverting or noninverting) buffer type. (ii) For simultaneous routing and buffering with a single noninverting buffer type, we give a factor 2(1+/spl epsiv/) approximation algorithm and prove that no algorithm can guarantee a factor smaller than 2 unless P=NP. For the case of a single inverting buffer type, we give a factor 4(1+/spl epsiv/) approximation algorithm. (iii) We give local-improvement and clustering based MBRP heuristics with improved practical performance, and present a comprehensive experimental study comparing the runtime/quality trade-offs of the proposed MBRP heuristics on test cases extracted from recent industrial designs. Charles J. Alpert, Andrew B. Kahng, Bao Liu 0001, Ion I. Mandoiu, Alex Zelikovsky |
ICCAD | 2 |
| 2001 | Buffered Steiner trees for difficult instancesabstractBuffer insertion has become an increasingly critical optimization in high performance design. The problem of finding a delay-optimal buffered Steiner tree has been an active area of research, and excellent solutions exist for most instances. However, current approaches fail to adequately solve a particular class of real-world “difficult” instances which are characterized by a large number of sinks, variations in sink criticalities, and varying polarity requirements. We propose a new Steiner tree construction called C-Tree for these instance types. When combined with van Ginneken style buffer insertion, C-Tree achieves higher quality solutions with fewer resources compared to traditional approaches. Charles J. Alpert, Milos Hrkic, Jiang Hu 0001, Andrew B. Kahng, John Lillis, Bao Liu 0001, Stephen T. Quay, Sachin S. Sapatnekar, A. J. Sullivan, Paul G. Villarrubia |
ISPD | 4 |
| 2001 | Practical Approximation Algorithms for Separable Packing Linear Programs
Feodor F. Dragan, Andrew B. Kahng, Ion I. Mandoiu, Sudhakar Muddu, Alex Zelikovsky |
WADS | 2 |
| 2001 | Limitations and challenges of computer-aided design technology for CMOS VLSIabstractAs manufacturing technology moves toward fundamental limits of silicon CMOS processing, the ability to reap the full potential of available transistors and interconnect is increasingly important. Design technology (DT) is concerned with the automated or semi-automated conception, synthesis, verification, and eventual testing of microelectronic systems. While manufacturing technology faces fundamental limits inherent in physical laws or material properties, design technology faces fundamental limitations inherent in the computational intractability of design optimizations and in the broad and unknown range of potential applications within various design processes. In this paper, we explore limitations to how design technology can enable the implementation of single-chip microelectronic systems that take full advantage of manufacturing technology with respect to such criteria as layout density performance, and power dissipation. Randal E. Bryant, Kwang-Ting Cheng, Andrew B. Kahng, Kurt Keutzer, Wojciech Maly, A. Richard Newton, Lawrence T. Pileggi, Jan M. Rabaey, Alberto L. Sangiovanni-Vincentelli |
Proc. IEEE | 3 |
| 2001 | Constraint-based watermarking techniques for design IP protectionabstractDigital system designs are the product of valuable effort and know-how. Their embodiments, from software and hardware description language program down to device-level netlist and mask data, represent carefully guarded intellectual property (IP). Hence, design methodologies based on IP reuse require new mechanisms to protect the rights of IP producers and owners. This paper establishes principles of watermarking-based IP protection, where a watermark is a mechanism for identification that is: (1) nearly invisible to human and machine inspection; (2) difficult to remove; and (3) permanently embedded as an integral part of the design. Watermarking addresses IP protection by tracing unauthorized reuse and making untraceable unauthorized reuse as difficult as recreating given pieces of IP from scratch. We survey related work in cryptography and design methodology, then develop desiderata, metrics, and concrete protocols for constraint-based watermarking at various stages of the very large scale integration (VLSI) design process. In particular, we propose a new preprocessing approach that embeds watermarks as constraints into the input of a black-box design tool and a new postprocessing approach that embeds watermarks as constraints into the output of a black-box design tool. To demonstrate that our protocols can be transparently integrated into existing design flows, we use a testbed of commercial tools for VLSI physical design and embed watermarks into real-world industrial designs. We show that the implementation overhead is low-both in terms of central processing unit time and such standard physical design metrics as wirelength, layout area, number of vias, and routing congestion. We empirically show that the placement and routing applications considered in our methods achieve strong proofs of authorship and are resistant to tampering and do not adversely influence timing. Andrew B. Kahng, John C. Lach, William H. Mangione-Smith, Stefanus Mantik, Igor L. Markov, Miodrag Potkonjak, Paul Tucker, Gregory Wolfe |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2001 | Toward accurate models of achievable routingabstractModels of achievable routing, i.e., chip wireability, rely on estimates of available and required routing resources. Required routing resources are estimated from placement or (a priori) using wire length estimation models. Available routing resources are estimated by calculating a nominal "supply" then take into account such factors as the efficiency of the router and the impact of vias. Models of achievable routing can be used to optimize interconnect process parameters for future designs or to supply objectives that guide layout tools to promising solutions. Such models must be accurate in order to be useful and must support empirical verification and calibration by actual routing results. In this paper, we discuss the validation of such models and we apply our validation process to three existing models. We find notable inaccuracies in the existing models when matched against real data. We then present a thorough analysis of the assumptions underlying these models. Based on this analysis, we discuss requirements for predictors of routing resources and make suggestions for a new model of achievable routing. Andrew B. Kahng, Stefanus Mantik, Dirk Stroobandt |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2000 | Improved algorithms for hypergraph bipartitioningabstractMultilevel Fiduccia-Mattheyses (MLFM) hypergraph partitioning [3,22,24] is a fundamental optimization in VLSI CAD physical design.The leading implementation, hMetis [23], has since 1997 pro veditself substantially superior in both runtime and solution quality t o e v en v ery recent w orks (e.g., [13,17,25]).In this work, we present t w o sets of results: (i) new techniques for at FM-based hypergraph partitioning (which is the core of multilev el implementations), and (ii) a new multilev el implementation that oers leadingedge performance.Our new techniques for at partitioning con rm the conjecture from [10], suggesting that specialized partitioning heuristics may be able to actively exploit xed nodes in partitioning instances arising in the driving top-down placement con text.Our FM varian t is competitive with traditional FM on instances without terminals [1] and considerably superior on instances with xed nodes (i.e., arising during top-down placement [8]).Our multilev el FM variant a v oids sev eral complex heuristics and non-trivial tunings that often lead to complex implementations; it ac hiev es trade-os bet w een solution quality and run time that are comparable or better than those achiev ed by hMetis-1.5.3 (the latest available version).F ollowing [6], we attempt to provide algorithm descriptions that are as detailed and unambiguous as possible, to allow replicabilit y and speed impro vements in future research. Andrew E. Caldwell, Andrew B. Kahng, Igor L. Markov |
ASP-DAC | 2 |
| 2000 | Monte-Carlo algorithms for layout density controlabstractAbstract| Chemical-mechanical polishing (CMP) and other manufacturing steps in very deep submicron VLSI have varying eects on devic e and inter connect features, dep ending on local char acteristics of the layout.T o enhance manufacturability and performance p r edictability, we seek to make the layout uniform with respect to prescribed densit ycriteria, by inserting \ ll" geometries into the layout.We propose several new Monte-Carlo based lling methods with fast dynamic data structures and report the tradeo between runtime and accuracy for the suggested methods.Compared to existing linear programming based a p p r oaches, our Monte-Carlo methods seem very promising as they produc enearly-optimal solutions within reasonable runtimes. Yu Chen 0005, Andrew B. Kahng, Gabriel Robins, Alex Zelikovsky |
ASP-DAC | 2 |
| 2000 | GTX: the MARCO GSRC technology extrapolation systemabstractTechnology extrapolation — the calibration and prediction of achievable design in future technology generations — drives the evolution of VLSI system architectures, design methodologies, and design tools. This paper describes initial experiences with development and use of GTX, the MARCO GSRC Technology Extrapolation system. GTX provides a robust, portable framework for interactive specification and comparison of modeling choices, e.g., for predicting system cycle time, die size and power dissipation. We use GTX to reveal surprising levels of uncertainty (modeling and parameter sensitivity) in widely-cited cycle-time models that drive recent roadmaps. We also describe new SOI and bulk device models that have been developed for GTX, as well as studies of power dissipation and delay uncertainty under various implementation assumptions for global interconnects. Andrew E. Caldwell, Yu Cao 0001, Andrew B. Kahng, Farinaz Koushanfar, Hua Lu 0004, Igor L. Markov, Michael Oliver, Dirk Stroobandt, Dennis Sylvester |
DAC | 3 |
| 2000 | Can recursive bisection alone produce routable placements?abstractThis work focuses on congestion-driven placement of standard cells into rows in the fixed-die context. We summarize the state-of-the-art after two decades of research in recursive bisection placement and implement a new placer, called Capo, to empirically study the achievable limits of the approach. From among recently proposed improvements to recursive bisection, Capo incorporates a leading-edge multilevel min-cut partitioner [7], techniques for partitioning with small tolerance [8], optimal min-cut partitioners and end-case min-wirelength placers [5], previously unpublished partitioning tolerance computations, and block splitting heuristics. On the other hand, our “good enough” implementation does not use “overlapping” [17], multi-way partitioners [17, 20], analytical placement, or congestion estimation [24, 35]. In order to run on recent industrial placement instances, Capo must take into account fixed macros, power stripes and rows with different allowed cell orientations. Capo reads industry-standard LEF/DEF, as well as formats of the GSRC bookshelf for VLSI CAD algorithms [6], to enable comparisons on available placement instances in the fixed-die regime. Andrew E. Caldwell, Andrew B. Kahng, Igor L. Markov |
DAC | 2 |
| 2000 | Practical iterated fill synthesis for CMP uniformityabstractWe propose practical iterated methods for layout density control for CMP uniformity, based on linear programming, Monte-Carlo and greedy algorithms. We experimentally study the tradeoffs between two main filling objectives: minimizing density variation, and minimizing the total amount of inserted fill. Comparisons with previous filling methods show the advantages of our new iterated Monte-Carlo and iterated greedy methods. We achieve near-optimal filling with respect to each of the objectives and for both density models (spatial density [3] and effective density [8]). Our new methods are more efficient in practice than linear programming [3] and more accurate than non-iterated Monte-Carlo approaches [1]. Yu Chen 0005, Andrew B. Kahng, Gabriel Robins, Alex Zelikovsky |
DAC | 2 |
| 2000 | METRICS: a system architecture for design process optimizationabstractWe describe METRICS, a system to recover design productivity via new infrastructure for design process optimization. METRICS seeks to treat system design and implementation as a science, rather than an art. A key precept is that measuring a design process is a prerequisite to optimizing it and continuously achieving maximum productivity. METRICS (i) unobtrusively gathers characteristics of design artifacts, design process, and communications during the system development effort, and (ii) analyzes and compares that data to analogous data from prior efforts. METRICS infrastructure consists of (i) a standard metrics schema, along with metrics transmittal capabilities embedded directly into EDA tools or into wrappers around tools; (ii) a metrics data warehouse and metrics reports; (iii) data mining and visualization capabilities for project prediction, tracking, and diagnosis. We give experiences and insights gained from development and deployment of METRICS within a leading SOC design flow. Stephen Fenstermaker, Andrew B. Kahng, Stefanus Mantik, Bart Thielges |
DAC | 3 |
| 2000 | On switch factor based analysis of coupled RC interconnectsabstractIn this paper we presen t a rank-one update method for updating reduced-order models of interconnect parasitics when driv e resistances or load capacitances change, as commonly occurs during timing analysis. These rank-one updates are extremely inexpensive, do not require reexamining the original in terconnect netw ork, and most importantly are provably equivalent to rereducing the original netw ork. This abstract contains the proof only for the case of varying the driver resistance, but examples are given to sho w that the exactness holds more generally. In particular, a cross-talk case is examined where the conductance matrix is singular. Andrew B. Kahng, Sudhakar Muddu, Egino Sarto |
DAC | 1 |
| 2000 | EDA meets.COM (panel session): how E-services will change the EDA business modelabstractThe skyrocketing complexities of SOC design have led to such design bottlenecks as iterations, poor scaling of tools, and inadequate computational power of desktop workstations and servers. At the same time, many organizations face long capital budgeting cycles for new hardware and software, and increased total cost of ownership (software, maintenance, integration, training, etc.) for their design capability. This panel addresses the “dot .com” phenomenon and its associated technology and infrastructure — e-services — which offer the promise of new design capabilities for chip designers and new business models for EDA companies. Issues include (1) for EDA, how will e-services transform sales channels, the push to interoperability, or how customer needs are addressed? (2) for ASIC vendors, will e-services change how end users design, and how vendors deploy their services? (3) for e-services providers and investors, what's real, what's the market, and how does this affect valuations and business strategy? Tom Quan, Andrew B. Kahng |
DAC | 3 |
| 2000 | Effects of Global Interconnect Optimizations on Performance Estimation of Deep Submicron DesignabstractIn this paper, we quantify the impact of global interconnect optimization techniques that address such design objectives as delay, peak noise, delay uncertainty due to noise, power, and cost. In doing so, we develop a new system-performance simulation model as a set of studies within the MARCO GSRC Technology Extrapolation (GTX) system. We model a typical point-to-point global interconnect and focus on accurate assessment of both circuit and design technology with respect to such issues as inductance, signal line shielding, dynamic delay, buffer placement uncertainty and repeater staggering. We demonstrate, for example, that optimal wire sizing models need to consider inductive effects-and that use of more accurate (-1,3) worst-case capacitive coupling noise switch factors substantially increases peak noise estimates compared to traditional (0,2) bounds. We also find that optimal repeater sizes are significantly smaller than conventional models would suggest, especially when considering energy-delay issues. Yu Cao 0001, Chenming Hu, Xuejue Huang, Andrew B. Kahng, Sudhakar Muddu, Dirk Stroobandt, Dennis Sylvester |
ICCAD | 4 |
| 2000 | Provably Good Global Buffering Using an Available Buffer Block PlanabstractTo implement high-performance global interconnect without impacting the performance of existing blocks, the use of buffer blocks is increasingly popular in structured-custom and block-based ASIC/SOC methodologies. Recent works by Cong et al. (1999) and Tang and Wong (2000) give algorithms to solve the buffer block planning problem. In this paper we address the problem of how to perform buffering of global nets given an existing buffer block plan. Assuming that global nets have been already decomposed into two-pin connections, we give a provably good algorithm based on a recent approach of Garg and Konemann (1998) and Fleischer (1999). Our method routes connections using available buffer blocks, such that required upper and lower bounds on buffer intervals-as well as wirelength upper bounds per connection-are satisfied. Our model allows more than one buffer to be inserted into any given connection. In addition, our algorithm observes buffer parity constraints, i.e., it will choose to use an inverter or a buffer (=co-located pair of inverters) according to source and destination signal parity. The algorithm outperforms previous approaches and has been validated on top-level layouts extracted from a recent high-end microprocessor design. Feodor F. Dragan, Andrew B. Kahng, Ion I. Mandoiu, Sudhakar Muddu, Alex Zelikovsky |
ICCAD | 2 |
| 2000 | On Mismatches between Incremental Optimizers and Instance Perturbations in Physical Design ToolsabstractThe incremental, "construct by correction" design methodology has become widespread in constraint-dominated DSM design. We study the problem of ECO for physical design domains in the general context of incremental optimization. We observe that an incremental design methodology is typically built from a full optimizer that generates a solution for an initial instance, and an incremental optimizer that generates a sequence of solutions corresponding to a sequence of perturbed instances. Our hypothesis is that in practice, there can be a mismatch between the strength of the incremental optimizer and the magnitude of the perturbation between successive instances. When such a mismatch occurs, the solution quality will degrade-perhaps to the point where the incremental optimizer should be replaced by the full optimizer. We document this phenomenon for three distinct domains-partitioning, placement and routing-using leading industry and academic tools. Our experiments show that current CAD tools may not be correctly designed for ECO-dominated design processes. Thus, compatibility between optimizer and instance perturbation merits attention both as a research question and as a matter of industry design practice. Andrew B. Kahng, Stefanus Mantik |
ICCAD | 1 |
| 2000 | Classical floorplanning harmful?abstractClassical floorplanning formulations may lead researchers to solve the wrong problems. This paper points out several examples, including (i) the preoccupation with packing-driven, as opposed to connectivity-driven, problem formulations and benchmarking standards; (ii) the preoccupation with rectangular (and L or T shaped) block shapes; and (iii) the lack of attention to algorithm scalability, fixed-die layout requirements, and the overall RTL-down methodology context. The right problem formulations must match the purpose and context of prevailing RTL-down design methodologies, and must be neither overconstrained nor underconstrained. The right solution ingredients are those which are scalable while delivering good solution quality according to relevant metrics. We also describe new problem formulations and solution ingredients, notably a perfect rectilinear floorplanning formulation that seeks zero-whitespace, perfectly packed rectilinear floorplans in a fixed-die regime. The paper closes with a list of questions for future research. Andrew B. Kahng |
ISPD | 1 |
| 2000 | Requirements for models of achievable routingabstractABSTRACT General Terms Models of achievable routing, i.e., chip wireability, rely on estimates of available and required routing resources. Re-quired routing resources are estimated from placement, or (a priori) using wirelength estimation models. Available rout-ing resources are estimated by calculating a nominal “sup-ply”, then taking into account such factors as the efficiency of the router and the impact of vias. Models of achievable routing can be used to optimize inter-connect process parameters for future designs or to supply objectives that guide layout tools to promising solutions. Such models must be accurate in order to be useful, and must support empirical verification and calibration by ac-tual routing results. In this paper, we discuss the validation of such models and we apply our validation process to three existing models. We find notable inaccuracies in the existing models when matched against real data. We then present a thorough analysis of the assumptions underlying these models; based on this analysis, we discuss requirements for predictors of routing resources within models of achievable routing. Andrew B. Kahng, Stefanus Mantik, Dirk Stroobandt |
ISPD | 1 |
| 2000 | Hypergraph partitioning with fixed vertices [VLSI CAD]abstractWe empirically assess the implications of fixed terminals for hypergraph partitioning heuristics. Our experimental testbed incorporates a leading-edge multilevel hypergraph partitioner and IBM-internal circuits that have recently been released as part of the ISPD-98 Benchmark Suite. We find that the presence of fixed terminals can make a partitioning instance considerably easier (possibly to the point of being "trivial"); much less effort is needed to stably reach solution qualities that are near best-achievable. Toward development of partitioning heuristics specific to the fixed-terminals regime, we study the pass statistics of flat FM-based partitioning heuristics. Our data suggest that more fixed terminals implies that the improvements within a pass will more likely occur near the beginning of the pass. Restricting the length of passes-which degrades solution quality in the classic (free-hypergraph) context-is relatively safe for the fixed-terminals regime and considerably reduces the run times of our FM-based heuristic implementations. The distinct nature of partitioning in the fixed-terminals regime has deep implications: (1) for the design and use of partitioners in top-down placement; (2) for the context in which VLSI hypergraph partitioning research is pursued; and (3) for the development of new benchmark instances for the research community. Charles J. Alpert, Andrew E. Caldwell, Andrew B. Kahng, Igor L. Markov |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2000 | Optimal phase conflict removal for layout of dark field alternatingphase shifting masksabstractWe describe new, efficient algorithms for layout modification and phase assignment for dark field alternating-type phase shifting masks in the single exposure regime. We make the following contributions. First, we suggest new two-coloring and compaction approach that simultaneously optimizes layout and phase assignment which is based on planar embedding of an associated conflict graph. We also describe additional approaches to cooptimization of layout and phase assignment for alternating PSM. Second, we give optimal and fast algorithms to minimize the number of phase conflicts that must be removed to ensure two colorability of the conflict graph. We reduce this problem to the T-join problem which asks for a minimum weight edge set A such that a node u is incident to an odd number of edges of A if u belongs to a given node subset T of a weighted graph. Third, we suggest several practical algorithms for the T-join problem. In sparse graphs, our algorithms are faster than previously known methods. Computational experience with industrial VLSI layout benchmarks shows the advantages of the new algorithms. Piotr Berman, Andrew B. Kahng, Devendra Vidhani, Alex Zelikovsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2000 | Optimal partitioners and end-case placers for standard-cell layoutabstractWe study alternatives to classic Fiduccia-Mattheyses (FM)-based partitioning algorithms in the context of end-case processing for top-down standard-cell placement. While the divide step in the top-down divide and conquer is usually performed heuristically, we observe that optimal solutions can be found for many sufficiently small partitioning instances. Our main motivation is that small partitioning instances frequently contain multiple cells that are larger than the prescribed partitioning tolerance, and that cannot be moved iteratively while preserving the legality of a solution. To sample the suboptimality of FM-based partitioning algorithms, we focus on optimal partitioning and placement algorithms based on either enumeration or branch-and-bound that are invoked for instances below prescribed size thresholds, e.g., <10 cells for placement and <30 cells for partitioning. Such partitioners transparently handle tight balance constraints and uneven cell sizes while typically achieving 40% smaller cuts than best of several FM starts for instances between ten and 50 movable nodes and average degree 2-3. Our branch-and-bound codes incorporate various efficiency improvements, using results for hypergraphs (1993) and a graph-specific algorithm (1996). We achieve considerable speed-ups over single FM starts on such instances on average. Enumeration-based partitioners relying on Gray codes, while easier to implement and taking less time for elementary operations, can only compete with branch-and-bound on very small instances, where optimal placers achieve reasonable performance as well. In the context of a top-down global placer, the right combination of optimal partitioners and placers can achieve up to an average of 10% wirelength reduction and 50% CPU time savings for a set of industry testcases. Our results show that run-time versus quality tradeoffs may be different for small problem instances than for common large benchmarks, resulting in different comparisons of optimization algorithms. We therefore suggest that alternative algorithms be considered and, as an example, present detailed comparisons with the flow-based balanced partitioner heuristic. Andrew E. Caldwell, Andrew B. Kahng, Igor L. Markov |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1999 | Design and Implementation of the Fiduccia-Mattheyses Heuristic for VLSI Netlist Partitioning
Andrew E. Caldwell, Andrew B. Kahng, Igor L. Markov |
ALENEX | 2 |
| 1999 | Function Smoothing with Applications to VLSI LayoutabstractWe present approximations to non-smooth continuous functions by differentiable functions which are parameterized by a scalar /spl beta/>0 and have convenient limit behavior as /spl beta//spl rarr/0. For standard numerical methods, this translates into a tradeoff between solution quality and speed. We show the utility of our approximations for wirelength and delay estimations used by analytical placers for VLSI layout. Our approximations lead to more "solvable" problems. Ross Baldick, Andrew B. Kahng, Andrew A. Kennings, Igor L. Markov |
ASP-DAC | 2 |
| 1999 | New Multilevel and Hierarchical Algorithms for Layout Density ControlabstractCertain manufacturing steps in very deep submicron VLSI involve chemical-mechanical polishing (CIMP) which has varying effects on device and interconnect features, depending on local layout characteristics. To reduce manufacturing variation due to CMP and to improve yield and performance predictability, the layout needs to be made uniform with respect to certain density criteria, by inserting "fill" geometries into the layout. This paper presents an efficient multilevel approach to density analysis that affords user-tunable accuracy. We also develop exact fill synthesis solutions based on combining multilevel analysis with a linear programming approach. Our methods apply to both flat and hierarchical designs. Andrew B. Kahng, Gabriel Robins, Anish Singh, Alex Zelikovsky |
ASP-DAC | 1 |
| 1999 | Optimization of Linear Placements for Wirelength Minimization with Free SitesabstractWe study a type of linear placement problem arising in detailed placement optimization of a given cell row in the presence of white-space (extra sites). In this single-row placement problem, the cell order is fixed within the row; all cells in other rows are also fixed. We give the first solutions to the single-row problem: (i) a dynamic programming technique with time complexity O(m/sup 2/) where m is the number of nets incident to cells in the given row, and (ii) an O(m log m) technique that exploits the convexity of the wirelength objective. We also propose an iterative heuristic for improving cell ordering within a row; this can be run optionally before applying either (i) or (ii). Experimental results show an average of 6.5% wirelength improvement on industry test cases when our methods are applied to the final output of a leading industry placement tool. Andrew B. Kahng, Paul Tucker, Alex Zelikovsky |
ASP-DAC | 1 |
| 1999 | Effective Iterative Techniques for Fingerprinting Design IPabstractWhile previous watermarking-based approaches to intellectual property protection (IPP) have asymmetrically emphasized the IP provider's rights, the true goal of IPP is to ensure the rights of both the IP provider and the IP buyer.Symmetric fingerprinting schemes have been widely and effectively used to achieve this goal; however, their application domain has been restricted only to static artifacts, such as image and audio.In this paper, we propose the first generic symmetric fingerprinting technique which can be applied to an arbitrary optimization/synthesis problem and, therefore, to hardware and software intellectual property.The key idea is to apply iterative optimization in an incremental fashion to solve a fingerprinted instance; this leverages the optimization effort already spent in obtaining a previous solution, yet generates a uniquely fingerprinted new solution.We use this approach as the basis for developing specific fingerprinting techniques for four important problems in VLSI CAD: partitioning, graph coloring, satisfiability, and standard-cell placement.We demonstrate the effectiveness of our fingerprinting techniques on a number of standard benchmarks for these tasks.Our approach provides an effective tradeoff between runtime and resilience against collusion. Andrew E. Caldwell, Hyun-Jin Choi, Andrew B. Kahng, Stefanus Mantik, Miodrag Potkonjak, Gang Qu 0001, Jennifer Wong-Ma |
DAC | 3 |
| 1999 | Hypergraph Partitioning for VLSI CAD: Methodology for Heuristic Development, Experimentation and ReportingabstractArticle Hypergraph partitioning for VLSI CAD: methodology for heuristic development, experimentation and reporting Share on Authors: Andrew E. Caldwell UCLA Computer Science Department, Los Angeles, CA UCLA Computer Science Department, Los Angeles, CAView Profile , Andrew B. Kahng UCLA Computer Science Department, Los Angeles, CA UCLA Computer Science Department, Los Angeles, CAView Profile , Andrew A. Kennings Cypress Semiconductor, Beaverton, OR Cypress Semiconductor, Beaverton, ORView Profile , Igor L. Markov UCLA Computer Science Department, Los Angeles, CA UCLA Computer Science Department, Los Angeles, CAView Profile Authors Info & Claims DAC '99: Proceedings of the 36th annual ACM/IEEE Design Automation ConferenceJune 1999 Pages 349–354https://doi.org/10.1145/309847.309955Online:01 June 1999Publication History 20citation418DownloadsMetricsTotal Citations20Total Downloads418Last 12 Months8Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Andrew E. Caldwell, Andrew B. Kahng, Andrew A. Kennings, Igor L. Markov |
DAC | 2 |
| 1999 | Hypergraph Partitioning with Fixed VerticesabstractWe empirically assess the implications of fixed terminals for hypergraph partitioning heuristics. Our experimental testbed incorporates a leading-edge multilevel hypergraph partitioner [14] [3] and IBM-internal circuits that have recently been released as part of the ISPD-98 Benchmark Suite [2, 1]. We find that the presence of fixed terminals can make a partitioning instance considerably easier (possibly to the point of being “trivial”): much less effort is needed to stably reach solution qualities that are near bestachievable. Toward development of partitioning heuristics specific to the fixed-terminals regime, we study the pass statistics of flat FM-based partitioning heuristics. Our data suggest that with more fixed terminals, the improvements in a pass are more likely to occur near the beginning of the pass. Restricting the length of passes – which degrades solution quality in the classic (free-hypergraph) context – is relatively safe for the fixed-terminals regime and considerably reduces run time of our FM-based heuristic implementations. We believe that the distinct nature of partitioning in the fixed-terminals regime has deep implications (i) for the design and use of partitioners in top-down placement, (ii) for the context in which VLSI hypergraph partitioning research is pursued, and (iii) for the development of new benchmark instances for the research community. Andrew E. Caldwell, Andrew B. Kahng, Igor L. Markov |
DAC | 2 |
| 1999 | Subwavelength Lithography and Its Potential Impact on Design and EDAabstractThis tutorial paper surveys the potential implications of subwavelength optical lithography for new tools and flows in the interface between layout design and manufacturability. We review control of optical process effects by optical proximity correction (OPC) and phase-shifting masks (PSM), then focus on the implications of OPC and PSM for layout synthesis and verification methodologies. Our discussion addresses the necessary changes in the designto-manufacturing flow, including infrastructure development in the mask and process communities, evolution of design methodology, and opportunities for research and development in the physical layout and verification areas of EDA. Andrew B. Kahng, Y. C. Pati |
DAC | 1 |
| 1999 | Subwavelength Lithography: How Will It Affect Your Design Flow? (Panel)abstractNo abstract available. Andrew B. Kahng, Y. C. Pati, Warren Grobman, Robert Pack, Lance A. Glasser |
DAC | 1 |
| 1999 | The associative-skew clock routing problemabstractWe introduce the associative skew clock routing problem, which seeks a clock routing tree such that zero skew is preserved only within identified groups of sinks. The associative skew problem is easier to address within current EDA frameworks than useful-skew (skew-scheduling) approaches, and defines an interesting tradeoff between the traditional zero-skew clock routing problem (one sink group) and the Steiner minimum tree problem (n sink groups). We present a set of heuristic building blocks, including an efficient and optimal method of merging two zero-skew trees such that zero skew is preserved within the sink sets of each tree. Finally, we list a number of open issues for research and practical application. Yu Chen 0005, Andrew B. Kahng, Gang Qu 0001, Alex Zelikovsky |
ICCAD | 2 |
| 1999 | Copy detection for intellectual property protection of VLSI designsabstractWe give the first study of copy detection techniques for VLSI CAD applications; these techniques are complementary to previous watermarking-based IP protection methods in finding and proving improper use of design IP. After reviewing related literature (notably in the text processing domain), we propose a generic methodology for copy detection based on determining basic elements within structural representations of solutions (IPs), calculating (context-independent) signatures for such elements, and performing fast comparisons to identify potential violators of IP rights. We give example implementations of this methodology in the domains of scheduling, graph coloring and gate-level layout; experimental results show the effectiveness of our copy detection schemes as well as the low overhead of implementation. We remark on open research areas, notably the potentially deep and complementary interaction between watermarking and copy detection. Andrew B. Kahng, Darko Kirovski, Stefanus Mantik, Miodrag Potkonjak, Jennifer Wong-Ma |
ICCAD | 1 |
| 1999 | Partitioning with terminals: a "new" problem and new benchmarksabstractThe presence of jixed terminals in hypergraph partitioning instances arising in top-down standard-cell placement makes such instances qualitatively different from the free hypergruphs that have driven the past two decades of VLSI CAD partitioning research.In this paper we empirically show that with fixed terminals in the instance, less effort is needed to stably reach a given solution quality.We then develop new benchmark formats that flexibly capture the presence of terminals and any geometric embedding information associated with the partitioning instance.Our new formats not only allow modeling of top-down placement, but also enable study of placement-specific partitioning objectives, e.g., based on net bounding boxes and Steiner tree estimators.Finally, we develop a new suite of partitioning benchmarks with fixed terminals, based on the actual placement data from the IBM-internal circuits released in the ISPD-98 Benchmark Suite [ 1, 21.A set of partitioning results is presented with nmtime regimes appropriate to the placement use model, as a baseline for future efforts in the research community. Charles J. Alpert, Andrew E. Caldwell, Andrew B. Kahng, Igor L. Markov |
ISPD | 3 |
| 1999 | Optimal phase conflict removal for layout of dark field alternating phase shifting masksabstractWe describe new, efficient algorithms for layout modification and phase assignment for dark field alternating-type phase shifting masks in the single exposure regime. We make the following contributions. First, we suggest new two-coloring and compaction approach that simultaneously optimizes layout and phase assignment which is based on planar embedding of an associated conflict graph. We also describe additional approaches to cooptimization of layout and phase assignment for alternating PSM. Second, we give optimal and fast algorithms to minimize the number of phase conflicts that must be removed to ensure two colorability of the conflict graph. We reduce this problem to the -join problem which asks for a minimum weight edge set such that a node is incident to an odd number of edges of if belongs to a given node subset of a weighted graph. Third, we suggest several practical algorithms for the -join problem. In sparse graphs, our algorithms are faster than previously known methods. Computational experience with industrial VLSI layout benchmarks shows the advantages of the new algorithms. Piotr Berman, Andrew B. Kahng, Devendra Vidhani, Alex Zelikovsky |
ISPD | 2 |
| 1999 | Optimal partitioners and end-case placers for standard-cell layoutabstractArticle Optimal partitioners and end-case placers for standard-cell layout Share on Authors: A. E. Caldwell UCLA Computer Science Dept., Los Angeles, CA UCLA Computer Science Dept., Los Angeles, CAView Profile , A. B. Kahng UCLA Computer Science Dept., Los Angeles, CA UCLA Computer Science Dept., Los Angeles, CAView Profile , I. L. Markov UCLA Computer Science Dept., Los Angeles, CA UCLA Computer Science Dept., Los Angeles, CAView Profile Authors Info & Claims ISPD '99: Proceedings of the 1999 international symposium on Physical designApril 1999 Pages 90–96https://doi.org/10.1145/299996.300032Published:12 April 1999 31citation350DownloadsMetricsTotal Citations31Total Downloads350Last 12 Months7Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Andrew E. Caldwell, Andrew B. Kahng, Igor L. Markov |
ISPD | 2 |
| 1999 | Subwavelength optical lithography: challenges and impact on physical designabstractWe review the implications of subwavelength optical lithography for new tools and flows in the interface between layout design and mtiufacturability.After discussing the necessity of corrections for optical process effects (i.e., use of optical proximity correction (OPC) and phase-shifting masks (PSM)), we focus on the implications of OPC and PSM for layout and verification methodologies.Our discussion addresses the necessary changes in the design-tomanufacturing 5ow, including infrastructure development in the mask and process communities as well as opportunities for research and development in physical layout and verification. Andrew B. Kahng, Y. C. Pati |
ISPD | 1 |
| 1999 | The T-join Problem in Sparse Graphs: Applications to Phase Assignment Problem in VLSI Mask Layout
Piotr Berman, Andrew B. Kahng, Devendra Vidhani, Alex Zelikovsky |
WADS | 2 |
| 1999 | Spectral Partitioning with Multiple Eigenvectors
Charles J. Alpert, Andrew B. Kahng, So-Zen Yao |
Discret. Appl. Math. | 2 |
| 1999 | On wirelength estimations for row-based placementabstractWirelength estimation in very large scale integration layout is fundamental to any predetailed routing estimate of timing or routability. In this paper, we develop efficient wirelength estimation techniques appropriate for wirelength estimation during top-down floorplanning and placement of cell-based designs. Our methods give accurate, linear-time approaches, typically with sublinear time complexity for dynamic updating of estimates (e.g., for annealing placement). Our techniques offer advantages not only for early on-line wirelength estimation during top-down placement, but also for a posteriori estimation of routed wirelength given a final placement. In developing these new estimators, we have made several contributions, including (1) insight into the contrast between region-based and bounding box-based rectilinear Steiner minimal tree (RStMT) estimation techniques; (2) empirical assessment of the correlations between pin placements of a multipin net that is contained in a block; and (3) new wirelength estimates that are functions of a block's complexity (number of cell instances) and aspect ratio. Andrew E. Caldwell, Andrew B. Kahng, Stefanus Mantik, Igor L. Markov, Alex Zelikovsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1999 | Filling algorithms and analyses for layout density controlabstractIn very deep-submicron very large scale integration (VLSI), manufacturing steps involving chemical-mechanical polishing (CMP) have varying effects on device and interconnect features, depending on local characteristics of the layout. To reduce manufacturing variation due to CMP and to improve performance predictability and yield, the layout must be made uniform with respect to certain density criteria, by inserting "fill" geometries into the layout. To date, only foundries and special mask data processing tools perform layout post-processing for density control. In the future, better convergence of performance verification flows will depend on such layout manipulations being embedded within the layout synthesis (place-and-route) flow. In this paper, we give the first realistic formulation of the filling problem that arises in layout optimization for manufacturability. Our formulation seeks to add features to a given process layer, such that (1) feature area densities satisfy prescribed upper and lower bounds in all windows of given size and (2) the maximum variation of such densities over all possible window positions in the layout is minimized. We present efficient algorithms for density analysis, notably a multilevel approach that affords user-tunable accuracy. We also develop exact solutions to the problem of fill synthesis, based on a linear programming approach. These include a linear programming (LP) formulation for the fixed-dissection regime (where density bounds are imposed on a predetermined set of windows in the layout) and an LP formulation that is automatically generated by our multilevel density analysis. We briefly review criteria for fill pattern synthesis, and the paper then concludes with computational results and directions for future research. Andrew B. Kahng, Gabriel Robins, Anish Singh, Alex Zelikovsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1998 | Watermarking Techniques for Intellectual Property ProtectionabstractDigital system designs are the product of valuable effort and knowhow. Their embodiments, from software and HDL program down to device-level netlist and mask data, represent carefully guarded intellectual property (IP). Hence, design methodologies based on IP reuse require new mechanisms to protect the rights of IP producers and owners. This paper establishes principles of watermarkingbased IP protection, where a watermark is a mechanism for identification that is (i) nearly invisible to human and machine inspection, (ii) difficult to remove, and (iii) permanently embedded as an integral part of the design. We survey related work in cryptography and design methodology, then develop desiderata, metrics and example approaches -- centering on constraint-based techniques -- for watermarking at various stages of the VLSI design process. 1 Introduction The advance of processing technology has led to a rapid increase in IC design complexity. The economic drivers are compelling: only by put... Andrew B. Kahng, John C. Lach, William H. Mangione-Smith, Stefanus Mantik, Igor L. Markov, Miodrag Potkonjak, Paul Tucker, Gregory Wolfe |
DAC | 1 |
| 1998 | Robust IP Watermarking Methodologies for Physical DesignabstractIncreasingly popular reuse-based design paradigms create a pressing need for authorship enforcement techniques that protect the intellectual property rights of designers. We develop the first intellectual property protection protocols for embedding design watermarks at the physical design level. We demonstrate that these protocols are tarnsparent with respect to existing industrial tools and design flows, and that they can embed watermarks into real-world industrial designs with very low implementation overhead (as measured by such standard metrics as wirelength, layout area, number of vias, routing congestion and CPU time). On several industrial test cases, we obtain extremely strong, tamper-resistant proofs of authorship for placement and routing solutions. Andrew B. Kahng, Stefanus Mantik, Igor L. Markov, Miodrag Potkonjak, Paul Tucker, Gregory Wolfe |
DAC | 1 |
| 1998 | Interconnect Tuning Strategies for High-Performance IcsabstractInterconnect tuning is an increasingly critical degree of freedom in the physical design of high-performance VLSI systems. By interconnect tuning, we refer to the selection of line thicknesses, widths and spacings in multi-layer interconnect to simultaneously optimize signal distribution, signal performance, signal integrity, and interconnect manufacturability and reliability. This is a key activity in most leading-edge design projects, but has received little attention in the literature. Our work provides the first technology-specific studies of interconnect tuning in the literature. We center on global wiring layers and interconnect tuning issues related to bus routing, repeater insertion, and choice of shielding/spacing rules for signal integrity and performance. We address four basic questions. (1) How should width and spacing be allocated to maximize performance for a given line pitch? (2) For a given line pitch, what criteria affect the optimal interval at which repeaters should be inserted into global interconnects? (3) Under what circumstances are shield wires the optimum technique for improving interconnect performance? (4) In global interconnect with repeaters, what other interconnect tuning is possible? Our study of question (4) demonstrates a new approach of offsetting repeater placements that can reduce worst-case cross-chip delays by over 30% in current technologies. Andrew B. Kahng, Sudhakar Muddu, Egino Sarto |
DATE | 1 |
| 1998 | On wirelength estimations for row-based placementabstractWirelength estimation in VLSI layout is fundamental to any pre-detailed routing estimate of timing or routability. In this paper, we develop new wirelength estimation techniques appropriate for top-down floor-planning and placement synthesis of row-based VLSI layouts. Our methods include accurate, linear-time approaches, often with sublinear time complexity for dynamic updating of estimates (e.g., for annealing placement). The new techniques offer advantages not only for early on-line wirelength estimation during top-down placement, but also for a posteriori estimation of routed wirelength given a final placement. In developing these new estimators, we have made several theoretical contributions. Notably, we have resolved the long-standing discrepancy between region-based and bounding box-based RSMT estimation techniques; this leads to new estimates that are functions of instance size n and aspect ratio AR. Andrew E. Caldwell, Andrew B. Kahng, Stefanus Mantik, Igor L. Markov, Alex Zelikovsky |
ISPD | 2 |
| 1998 | Futures for partitioning in physical design (tutorial)abstractThe context for partitioning in physical design is dominated by two concerns: top-down design and the focus on spatial embedding. The role of partitioning is exactly that of a facilitator of divide-and-conquer metaheuristics for floorplanning, timing and placement optimization. Formulations or optimization objectives for partitioning follow from its context and role. Finally, the available algorithm technology determines how effectively we can address a given partitioning formulation and optimize a given objective. This invited paper considers the future of partitioning for physical design in light of these factors, and proposes a list of technology needs. A living version of this paper can be found at vlsicad.cs.ucla.edu. Andrew B. Kahng |
ISPD | 1 |
| 1998 | New efficient algorithms for computing effective capacitanceabstractWe describe a novel iterationless approach for computing the effective capacitance of an interconnect load at a driving gate output. Our new approach is considerably faster than previous methods for computing effective capacitance, with little or no loss of accuracy. Thus, the approach is suitable within the analysis loop for performance-driven iterative layout optimization. After reviewing previous gate load models and effective capacitance approximations, we separately derive our method for the cases of step and ramp waveform at the gate output, and note on-going extensions for the case of complex gates (e.g., channel-connected components). Experimental results using the new effective capacitance approach show that our resulting delay estimates are quite accurate — within 15% of IISPICE-computed delays on data corresponding to an 0.25µm microprocessor design. Andrew B. Kahng, Sudhakar Muddu |
ISPD | 1 |
| 1998 | Filling and slotting: analysis and algorithmsabstractIn very deep-submicron VLSI, certain manufacturing steps &mdash notably optical exposure, resist development and etch, chemical vapor deposition and chemical-mechanical polishing (CMP)&mdash have varying effects on device and interconnect features depending on local characteristics of the layout. To make these effects uniform and predictable, the layout itself must be made uniform with respect to certain density parameters. Traditionally, only foundries have performed the post-processing needed to achieve this uniformity, via insertion (“filling”) or partial deletion (“slotting”) of features in the layout. Today, however, physical design and verification tools cannot remain oblivious to such foundry post-processing. Without an accurate estimate of the filling and slotting, RC extraction, delay calculation, and timing and noise analysis flows will all suffer from wild inaccuracies. Therefore, future place-and-route tools must efficiently perform filling and slotting prior to performance analysis within the layout optimization loop. We give the first formulations of the filling and slotting problems that arise in layout post-processing or layout optimization for manufacturability. Such formulations seek to add or remove features to a given process layer, so that the local area or perimeter density of features satisfies prescribed upper and lower bounds in all windows of a given size. We also present efficient algorithms for density analysis as well as for filling/slotting synthesis. Our work provides a new unification between manufacturing and physical design, and captures a number of general requirements imposed on layout by the manufacturing process. Andrew B. Kahng, Gabriel Robins, Anish Singh, Alex Zelikovsky |
ISPD | 1 |
| 1998 | How to test a treeabstractWe address the problem of verifying that a tree is connected using probe operations which check mutual connectivity between two (or more) leaves of the tree. We present optimal algorithms for determining minimal probe sets that detect all possible edge and vertex faults in arbitrary trees. Our results are of particular interest for the testing of interconnection substrates in VLSI multichip module packaging technologies. © 1998 John Wiley & Sons, Inc. Networks 32: 189–197, 1998 Andrew B. Kahng, Gabriel Robins, Elizabeth A. Walkup |
Networks | 1 |
| 1998 | Faster minimization of linear wirelength for global placementabstractA linear wirelength objective more effectively captures timing, congestion, and other global placement considerations than a squared wirelength objective. The GORDIAN-L cell placement tool minimizes linear wirelength by first approximating the linear wirelength objective by a modified squared wirelength objective, then executing the following loop-(1) minimize the current objective to yield some approximate solution and (2) use the resulting solution to construct a more accurate objective-until the solution converges. This paper shows how to apply a generalization of an algorithm due to Weiszfeld (1937) to placement with a linear wirelength objective and that the main GORDIAN-L loop is actually a special case of this algorithm. We then propose applying a regularization parameter to the generalized Weiszfeld algorithm to control the tradeoff between convergence and solution accuracy; the GORDIAN-L iteration is equivalent to setting this regularization parameter to zero. We also apply novel numerical methods, such as the primal-Newton and primal-dual Newton iterations, to optimize the linear wirelength objective. Finally, we show both theoretically and empirically that the primal-dual Newton iteration stably attains quadratic convergence, while the generalized Weiszfeld iteration is linear convergent. Hence, primal-dual Newton is a superior choice for implementing a placer such as GORDIAN-L, or for any linear wirelength optimization. Charles J. Alpert, Tony F. Chan, Andrew B. Kahng, Igor L. Markov, Pep Mulet |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1998 | Multilevel circuit partitioningabstractMany previous works in partitioning have used some underlying clustering algorithm to improve performance. As problem sizes reach new levels of complexity, a single application of a clustering algorithm is insufficient to produce excellent solutions. Recent work has illustrated the promise of multilevel approaches. A multilevel partitioning algorithm recursively clusters the instance until its size is smaller than a given threshold, then unclusters the instance, while applying a partitioning refinement algorithm. In this paper, we propose a new multilevel partitioning algorithm that exploits some of the latest innovations of classical iterative partitioning approaches. Our method also uses a new technique to control the number of levels in our matching-based clustering algorithm. Experimental results show that our heuristic outperforms numerous existing bipartitioning heuristics with improvements ranging from 6.9 to 27.9% for 100 runs and 3.0 to 20.6% for just ten runs (while also using less CPU time). Further, our algorithm generates solutions better than the best known mincut bipartitionings for seven of the ACM/SIGDA benchmark circuits, including golem3 (which has over 100000 cells). We also present quadrisection results which compare favorably to the partitionings obtained by the GORDIAN cell placement tool. Our work in multilevel quadrisection has been used as the basis for an effective cell placement package. Charles J. Alpert, Dennis J.-H. Huang, Andrew B. Kahng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1998 | Efficient algorithms for the minimum shortest path Steiner arborescence problem with applications to VLSI physical designabstractGiven an undirected graph G=(V, E) with positive edge weights (lengths) /spl omega/: E/spl rarr//spl Rfr//sup +/, a set of terminals (sinks) N/spl sube/V, and a unique root node e/spl epsiv/N, a shortest path Steiner arborescence (hereafter an arborescence) is a Steiner tree rooted at, spanning all terminals in N such that every source-to-sink path is a shortest path in G. Given a triple (G, N, r), the minimum shortest path Steiner arborescence (MSPSA) problem seeks an arborescence with minimum weight. The MSPSA problem has various applications in the areas of physical design of very large-scale integrated circuits (VLSI), multicast network communication, and supercomputer message routing; various eases have been studied in the literature. In this paper, we propose several heuristics and exact algorithms for the MSPSA problem with applications to VLSI physical design. Experiments indicate that our heuristics generate near optimal results and achieve speedups of orders of magnitude over existing algorithms. Jason Cong, Andrew B. Kahng, Kwok-Shing Leung |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1998 | Guest Editorial
Andrew B. Kahng, Majid Sarrafzadeh |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1998 | Bounded-skew clock and Steiner routingabstractWe study the minimum-cost bounded-skew routing tree problem under the pathlength (linear) and Elmore delay models. This problem captures several engineering tradeoffs in the design of routing topologies with controlled skew. Our bounded-skew routing algorithm, called the BST/DME algorithm, extends the DME algorithm for exact zero-skew trees via the concept of a merging region . For a prescribed topology , BST/DME constructs a bounded-skew tree (BST) in two phases: (i) a bottom-up phase to construct a binary tree of merging regions which represent the loci of possible embedding points of the internal nodes, and (ii) a top-down phase to determine the exact locations of the internal nodes. We present two approaches to construct the merging regions: (i) the Boundary Merging and Embedding (BME) method which utilizes merging points that are restricted to the boundaries of merging regions, and (ii) the Interior Merging and Embedding (IME) algorithm which employs a sampling strategy and a dynamic programming-based selection technique to consider merging points that are interior to, as well as on the boundary of, the merging regions. When the topology is not prescribed, we propose a new Greedy -BST/DME algorithm which combines the merging region computation with topology generation. The Greedy-BST/DME algorithm very closely matches the best known heuristics for the zero-skew case and for the unbounded-skew case (i.e., the Steiner minimal tree problem). Experimental results show that our BST algorithms can produce a set of routing solutions with smooth skew and wire length tradeoffs. Jason Cong, Andrew B. Kahng, Cheng-Kok Koh, Chung-Wen Albert Tsao |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 1997 | Multilevel Circuit PartitioningabstractRecent work has illustrated the promise ofmultilevel approaches for partitioning large circuits. Multilevel partitioningrecursively clusters the instance until its size is smallerthan a given threshold, then unclusters the instance while applyinga partitioning refinement algorithm. Our multilevel partitioner usesa new technique to control the number of levels in the matching-basedclustering phase and also exploits recent innovations in classiciterative partitioning. Our heuristic outperforms numerousexisting bipartitioning heuristics, with improvements rangingfrom 6.9 to 27.9% for 100 runs and 3.0 to 20.6% for just 10 runs(while also using less CPU time). Charles J. Alpert, Dennis J.-H. Huang, Andrew B. Kahng |
DAC | 3 |
| 1997 | Analysis and Justification of a Simple, Practical 2 1/2-D Capacitance Extraction MethodologyabstractThis paper addresses post-routing capacitance extraction duringperformance-driven layout. We first show how basic driversin process technology (planarization and minimum metal densityrequirements) actually simplify the extraction problem; wedo this by proposing and validating five "foundations" throughdetailed experiments with representative 0.18μm process parametersand a 3-D field solver.We then present a simple yetaccurate 2 1/2-D extraction methodology directly based on thefoundations.This methodology has been productized and isbeing shipped with the Cadence Silicon Ensemble 5.0 product.We conclude that the 2 1/2-D approach has sufficient accuracyfor current and near-term process generations. Jason Cong, Lei He 0001, Andrew B. Kahng, David Noice, Nagesh Shirali, Steve H.-C. Yen |
DAC | 3 |
| 1997 | More Practical Bounded-Skew Clock RoutingabstractAcademic clock routing research results has often hadlimited impact on industry practice, since such practical considerationsas hierarchical buffering, rise-time and overshoot constraints,obstacle- and legal location-checking, varying layer parasitics andcongestion, and even the underlying design flow are often ignored.This paper explores directions in which traditional formulationscan be extended so that the resulting algorithms are more usefulin production design environments. Specifically, the following issuesare addressed: (i) clock routing for varying layer parasiticswith nonzero via parasitics; (ii) obstacle-avoidance clock routing;(iii) a new topology design rule for prescribed-delay clock routing;and (iv) predictive modeling of the clock routing itself. Wedevelop new theoretical analyses and heuristics, and present experimentalresults that validate our new approaches. Andrew B. Kahng, Chung-Wen Albert Tsao |
DAC | 1 |
| 1997 | Faster minimization of linear wirelength for global placementabstractArticle Faster minimization of linear wirelength for global placement Share on Authors: Charles J. Alpert IBM Austin Research Laboratory, Austin, TX IBM Austin Research Laboratory, Austin, TXView Profile , Tony F. Chan UCLA Mathematics Department, Los Angeles, CA UCLA Mathematics Department, Los Angeles, CAView Profile , Dennis J.-H. Huang UCLA Computer Science Department, Los Angeles, CA UCLA Computer Science Department, Los Angeles, CAView Profile , Andrew B. Kahng UCLA Computer Science Department, Los Angeles, CA and Cadence Design Systems, Inc., San Jose, CA UCLA Computer Science Department, Los Angeles, CA and Cadence Design Systems, Inc., San Jose, CAView Profile , Igor L. Markov UCLA Mathematics Department, Los Angeles, CA UCLA Mathematics Department, Los Angeles, CAView Profile , Pep Mulet UCLA Mathematics Department, Los Angeles, CA UCLA Mathematics Department, Los Angeles, CAView Profile , Kenneth Yan UCLA Computer Science Department, Los Angeles, CA UCLA Computer Science Department, Los Angeles, CAView Profile Authors Info & Claims ISPD '97: Proceedings of the 1997 international symposium on Physical designApril 1997 Pages 4–11https://doi.org/10.1145/267665.267670Online:01 April 1997Publication History 17citation150DownloadsMetricsTotal Citations17Total Downloads150Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Charles J. Alpert, Tony F. Chan, Dennis J.-H. Huang, Andrew B. Kahng, Igor L. Markov, Pep Mulet, Kenneth Yan |
ISPD | 4 |
| 1997 | Efficient heuristics for the minimum shortest path Steiner arborescence problem with applications to VLSI physical designabstractGiven an undirected graph G = (K E) with positive edge weights (lengths) w : E + R+, a set of terminals (sinks) N C V, and a unique root node r E N, a shortest-path Ste%er arborescence (simply called an arborescence in the following) is a Steiner tree rooted at T spanning all terminals in N such that every source-to-sink path is a shortest path in G. Given a triple (G, N, T), the Minimum Shortest-Path Steiner Arborescence (MSPSA) problem seeks an arborescence with minimum weight.The MSPSA problem has various applications in the areas of VLSI physical design, multicast network communication, and supercomputer message routing; various cases have been studied in the literature.In this paper, we propose three efficient heuristics for the MSPSA problem and present applications to VLSI physical design.Experiments indicate that our algorithms generate near-optimal results and achieve speedups of several orders of magnitude over existing algorithms.1. Jason Cong, Andrew B. Kahng, Kwok-Shing Leung |
ISPD | 2 |
| 1997 | Partitioning-based standard-cell global placement with an exact objectiveabstractWe present a nexv top-down quadrisection-based global placer for standard-cell layout.The key contribution is a new general gain update scheme for partitioning that can exactly capture detailed placement objectives on a per-net basis.We use this gain update scheme, along with an efficient multilevel partitioner, as the basis for a new quadrisection-based placer called QUAD.Even though QUAD is a global placer, it can achieve significant improve ments in wirelength and congestion distribution over GOR-DIAN-L/DOMINO[SDJSI] [DJS94] (a leading quadratic placer with linear wirelength objective and detailed placement improvement).QUAD can be easily extended to capture various practical considerations; our timing-driven placement can obtain wirelength savings (as well as small cycle time improvements) versus the SPEED [RE95].1. Dennis J.-H. Huang, Andrew B. Kahng |
ISPD | 2 |
| 1997 | On implementation choices for iterative improvement partitioning algorithmsabstractIterative improvement partitioning algorithms such as the FM algorithm of Fiduccia and Mattheyses (1982), the algorithm of Krishnamurthy (1984), and Sanchis's extensions of these algorithms to multiway partitioning (1989) all rely on efficient data structures to select the modules to be moved from one partition to the other. The implementation choices for one of these data structures, the gain bucket, is investigated. Surprisingly, selection from gain buckets maintained as last-in-first-out (LIFO) stacks leads to significantly better results than gain buckets maintained randomly (as in previous studies of the FM algorithm or as first-in-first-out (FIFO) queues. In particular, LIFO buckets result in a 36% improvement over random buckets and a 43% improvement over FIFO buckets for minimum-cut bisection. Eliminating randomization from the bucket selection not only improves the solution quality, but has a greater impact on FM performance than adding the Krishnamurthy gain vector. The LIFO selection scheme also results in improvement over random schemes for multiway partitioning and for more sophisticated partitioning strategies such as the two-phase FM methodology. Finally, by combining insights from the LIFO gain buckets with the Krishnamurthy higher-level gain formulation, a new higher-level gain formulation is proposed. This alternative formulation results in a further 22% reduction in the average cut cost when compared directly to the Krishnamurthy formulation for higher-level gains, assuming LIFO organization for the gain buckets. Lars W. Hagen, Dennis J.-H. Huang, Andrew B. Kahng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1997 | Combining problem reduction and adaptive multistart: a new technique for superior iterative partitioningabstractVLSI netlist partitioning has been addressed chiefly by iterative methods (e.g. Kernighan-Lin and Fiduccia-Mattheyses) and spectral methods (e.g. Hagen-Kahng). Iterative methods are the de facto industry standard, but suffer diminished stability and solution quality when instances grow large. Spectral methods have achieved high-quality solutions, particularly for the ratio cut objective, but suffer excessive memory requirements and the inability to capture practical constraints (preplacements, variable module areas, etc.). This work develops a new approach to Fiduccia-Mattheyses (FM)-based iterative partitioning. We combine two concepts: (1) problem reduction using clustering and the two-phase FM methodology and (2) adaptive multistart, i.e. the intelligent selection of starting points for the iterative optimization, based on the results of previous optimizations. The resulting clustered adaptive multistart (CAMS) methodology substantially improves upon previous partitioning results in the literature, for both unit module areas and actual module areas, and for both the min-cut bisection and minimum ratio cut objectives. The CAMS method is surprisingly fast and has very stable solution quality, even for large benchmark instances. It has been applied as the basis of a clustering methodology within an industry placement tool. Lars W. Hagen, Andrew B. Kahng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1997 | An analytical delay model for RLC interconnectsabstractElmore delay has been widely used to estimate interconnect delays in the performance-driven synthesis and layout of very-large-scale-integration (VLSI) routing topologies. For typical RLC interconnections, however, Elmore delay can deviate significantly from SPICE-computed delay, since it is independent of inductance of the interconnect and rise time of the input signal. Here, we develop an analytical delay model based on first and second moments to incorporate inductance effects into the delay estimate for interconnection lines under step input. Delay estimates using our analytical model are within 15% of SPICE-computed delay across a wide range of interconnect parameter values. We also extend our delay model for estimation of source-sink delays in arbitrary interconnect trees. We observe significant improvement in the accuracy of delay estimates for interconnect trees when compared to the Elmore model, yet our estimates are as easy to compute as Elmore delay. Evaluation of our analytical models is several orders of magnitude faster than simulation using SPICE. We also illustrate the application of our model in controlling response undershoot/overshoot and reducing interconnect delay through constraints on the moments. Andrew B. Kahng, Sudhakar Muddu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1997 | Analysis of RC interconnections under ramp inputabstractWe give new methods for calculating the time-domain response for a finite-length distributed RC line that is stimulated by a ramp input. The following are our contributions. First, we obtain the solution of the diffusion equation for a seminfinite distributed RC line with ramp input. We then present a general and, in the limit, exact approach to compute the time-domain response for finite-length RC lines under ramp input by summing distinct diffusions starting at either end of the line. Next, we obtain analytical expressions for the finite time-domain voltage response for an open-ended finite RC line and for a finite RC line with capacitive load. The delay estimates using this method are very close to SPICE-computing delays. Finally, we present a general recursive equation for computing the higher-order diffusion components due to reflections at the source and load ends. Future work extends our method to response computations in general interconnection trees by modeling both reflection and transmission coefficients at discontinuities. Andrew B. Kahng, Sudhakar Muddu |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 1996 | Analysis of RC Interconnections Under Ramp InputabstractArticle Analysis of RC interconnections under ramp input Share on Authors: Andrew B. Kahng UCLA Computer Science Department, Los Angeles, CA UCLA Computer Science Department, Los Angeles, CAView Profile , Sudhakar Muddu UCLA Computer Science Department, Los Angeles, CA UCLA Computer Science Department, Los Angeles, CAView Profile Authors Info & Claims DAC '96: Proceedings of the 33rd annual Design Automation ConferenceJune 1996 Pages 533–538https://doi.org/10.1145/240518.240619Online:01 June 1996Publication History 6citation214DownloadsMetricsTotal Citations6Total Downloads214Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Andrew B. Kahng, Sudhakar Muddu |
DAC | 1 |
| 1996 | Analytical delay models for VLSI interconnects under ramp inputabstractElmore delay has been widely used as an analytical estimate of interconnect delays in the performance-driven synthesis and layout of VLSI routing topologies. However, for typical RLC interconnections with ramp input, Elmore delay can deviate by up to 100% or more from SPICE-computed delay since it is independent of rise time of the input ramp signal. We develop new analytical delay models based on the first and second moments of the interconnect transfer function when the input is a ramp signal with finite rise time. Delay estimates using our first moment based analytical models are within 4% of SPICE-computed delay, and models based on both first and second moments are within 2.3% of SPICE, across a wide range of interconnect parameter values. Evaluation of our analytical models is several orders of magnitude faster than simulation using SPICE. We also describe extensions of our approach for estimation of source-sink delays in arbitrary interconnect trees. Andrew B. Kahng, Kei Masuko, Sudhakar Muddu |
ICCAD | 1 |
| 1996 | Planar-DME: a single-layer zero-skew clock tree routerabstractThis paper presents new single-layer, i.e., planar-embeddable, clock tree constructions with exact zero skew under either the linear or the Elmore delay model. Our method, called Planar-DME, consists of two parts. The first algorithm, called Linear-Planar-DME, guarantees an optimal planar zero-skew clock tree (ZST) under the linear delay model. The second algorithm, called Elmore-Planar-DME, uses the Linear-Planar-DME connection topology in constructing a low-cost ZST according to the Elmore delay model. While a planar ZST under the linear delay model is easily converted to a planar ZST under the Elmore model by elongating tree edges in bottom-up order, our key idea is to avoid unneeded wire elongation by iterating the DME construction of ZST and the bottom-up modification of the resulting nonplanar routing. Costs of our planar ZST solutions are comparable to those of the best previous nonplanar ZST solutions, and substantially improve over previous planar clock routing methods. Andrew B. Kahng, Chung-Wen Albert Tsao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1996 | A general framework for vertex orderings with applications to circuit clusteringabstractVertex orderings have been successfully applied to problems in netlist clustering and for system partitioning and layout. We present a vertex ordering construction that encompasses most reasonable graph traversals. Two parameters-an attraction function and a window-provide the means for achieving various graph traversals and addressing particular clustering requirements. We then use dynamic programming to optimality split the vertex ordering into a multiway clustering. Our approach outperforms several clustering methods in the literature in terms of three distinct clustering objectives. The ordering construction, by itself, also outperforms existing graph ordering constructions for this application. Tuning our approach to "meta-objectives", particularly clustering for two-phase Fiduccia-Mattheyses bipartitioning, remains an open area of research. Charles J. Alpert, Andrew B. Kahng |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1995 | Quantified Suboptimality of VLSI Layout HeuristicsabstractWe show h o w t o quantify the suboptimality o f heuristic algorithms for NP-hard problems arising in VLSI layout.Our approach is based on the notion of constructing new scaled instances from an initial problem instance.From the given problem instance, we essentially construct doubled, tripled, etc. instances which h a v e optimum solution costs at most twice, three times, etc. that of the original instance.By executing the heuristic on these scaled instances, and then comparing the growth of solution cost with the growth of instance size, we can measure the scaling suboptimality of the heuristic.We give experimentally determined scaling behavior of several placement and partitioning heuristics; these results suggest that sigini cant improvement remains possible over current state-of-the-art methods. Lars W. Hagen, Dennis J.-H. Huang, Andrew B. Kahng |
DAC | 3 |
| 1995 | On the Bounded-Skew Clock and Steiner Routing ProblemsabstractWe study the minimum-cost bounded-skewrouting tree (BST) problem under the linear delay model.This problem captures several engineering tradeoffs in the design of routing topologies with controlled skew.We propose three tradeoff heuristics.(1) For a fixed topology Extended-DME (Ex-DME) extends the DME algorithm for exact zero-skew trees via the concept of a merging region.(2) For arbitrary topology and arbitrary embedding, Extended Greedy-DME (ExG-DME) very closely matches the best known heuristics for the zero-skew case,and for the infinite-skew case (i.e., the Steiner minimal tree problem).(3) For arbitrary topology and single-layer (planar) embedding, the Extended Planar-DME (ExP-DME) algorithm exactly matches the best known heuristic for zero-skew planar routing, and closely approaches the best known performance for the infinite-skew case.Our work provides unifications of the clock routing and Steiner tree heuristic literatures and gives smooth cost-skew tradeoff that enable good engineering solutions. Dennis J.-H. Huang, Andrew B. Kahng, Chung-Wen Albert Tsao |
DAC | 2 |
| 1995 | Multi-way System Partitioning into a Single Type or Multiple Types of FPGAsabstractThis paper considers the problem of partitioning a circuit into a collection of subcircuits, such that each subcircuit is feasible for some device from an FPGA library, and the total cost of devices is minimized. We propose a three-phase heuristic that uses ordering, clustering, and dynamic programming to achieve good solutions. Experimental comparisons are made with the previous methods of [4][9]. Dennis J.-H. Huang, Andrew B. Kahng |
FPGA | 2 |
| 1995 | Bounded-skew clock and Steiner routing under Elmore delayabstractWe study the minimum-cost bounded-skew routing tree problem under the Elmore delay model. We present two approaches to construct bounded-skew routing trees: (i) the Boundary Merging and Embedding (BME) method which utilizes merging points that are restricted to the boundaries of merging regions, and (ii) the Interior Merging and Embedding (IME) algorithm which employs a sampling strategy and dynamic programming to consider merging points that are interior to, rather than on the boundary of, the merging regions. Our new algorithms allow accurate control of Elmore delay skew, and show the utility of merging points inside merging regions. Jason Cong, Andrew B. Kahng, Cheng-Kok Koh, Chung-Wen Albert Tsao |
ICCAD | 2 |
| 1995 | Cooperative mobile robotics: antecedents and directionsabstractThere has been increased research interest in systems composed of multiple autonomous mobile robots exhibiting collective behavior. Groups of mobile robots are constructed, with an aim to studying such issues as group architecture, resource conflict, origin of cooperation, learning, and geometric problems. As yet, few applications of collective robotics have been reported, and supporting theory is still in its formative stages. In this paper, the authors give a critical survey of existing works and discuss open problems in this field, emphasizing the various theoretical issues that arise in the study of cooperative robotics. The authors describe the intellectual heritages that have guided early research, as well as possible additions to the set of existing motivations. Y. Uny Cao, Alex S. Fukunaga, Andrew B. Kahng, F. Meng |
IROS (1) | 3 |
| 1995 | Old Bachelor Acceptance: A New Class of Non-Monotone Threshold Accepting MethodsabstractStochastic hill-climbing algorithms, particularly simulated annealing (SA) and threshold acceptance (TA), have become very popular for global optimization applications. Typical implementations of SA or TA use monotone temperature or threshold schedules, and are not formulated to accommodate practical time limits. We present a new threshold acceptance strategy called Old Bachelor Acceptance (OBA), which has three distinguishing features: (i) it is specifically motivated by the practical requirement of optimization within a prescribed time bound, (ii) the threshold schedule is self-tuning, and (iii) the threshold schedule is non-monotone, with threshold values even allowed to become negative. The standard implementation of the TA method of Dueck and Scheuer is a special case of OBA. Experiments using several classes of symmetric traveling salesman problem instances show that OBA can outperform previous hill-climbing methods for time-critical optimizations. A number of directions for future work are suggested. INFORMS Journal on Computing, ISSN 1091-9856, was published as ORSA Journal on Computing from 1989 to 1995 under ISSN 0899-1499. T. C. Hu, Andrew B. Kahng, Chung-Wen Albert Tsao |
INFORMS J. Comput. | 2 |
| 1995 | Recent directions in netlist partitioning: a survey
Charles J. Alpert, Andrew B. Kahng |
Integr. | 2 |
| 1995 | Prim-Dijkstra tradeoffs for improved performance-driven routing tree designabstractAnalysis of Elmore delay in distributed RC tree structures shows the influence of both tree cost and tree radius on signal delay in VLSI interconnects. We give new and efficient interconnection tree constructions that smoothly combine the minimum cost and the minimum radius objectives, by combining respectively optimal algorithms due to Prim (1957) and Dijkstra (1959). Previous "shallow-light" techniques are both less direct and less effective: in practice, our methods achieve uniformly superior cost-radius tradeoffs. Timing simulations for a range of IC and MCM interconnect technologies show that our wirelength savings yield reduced signal delays when compared to shallow-light or standard minimum spanning tree and Steiner tree routing.> Charles J. Alpert, T. C. Hu, Dennis J.-H. Huang, Andrew B. Kahng, David R. Karger |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 1995 | Multiway partitioning via geometric embeddings, orderings, and dynamic programmingabstractThis paper presents effective algorithms for multiway partitioning. Confirming ideas originally due to Hall (1970), we demonstrate that geometric embeddings of the circuit netlist can lead to high-quality k-way partitionings. The netlist embeddings are derived via the computation of d eigenvectors of the Laplacian for a graph representation of the netlist. As Hall did not specify how to partition such geometric embeddings, we explore various geometric partitioning objectives and algorithms, and find that they are limited because they do not integrate topological information from the netlist. Thus, we also present a new partitioning algorithm that exploits both the geometric embedding and netlist information, as well as a restricted partitioning formulation that requires each cluster of the k-way partitioning to be contiguous in a given linear ordering. We begin with a d-dimensional spectral embedding and construct a heuristic 1-dimensional ordering of the modules (combining spacefilling curve with 3-Opt approaches originally proposed for the traveling salesman problem). We then apply dynamic programming to efficiently compute the optimal k-way split of the ordering for a variety of objective functions, including Scaled Cost and Absorption. This approach can transparently integrate user-specified cluster size bounds. Experiments show that this technique yields multiway partitionings with lower Sealed Cost than previous spectral approaches.> Charles J. Alpert, Andrew B. Kahng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1995 | Near-optimal critical sink routing tree constructionsabstractWe present critical-sink routing tree (CSRT) constructions which exploit available critical-path information to yield high-performance routing trees. Our CS-Steiner and "global slack removal" algorithms together modify traditional Steiner tree constructions to optimize signal delay at identified critical sinks. We further propose an iterative Elmore routing tree (ERT) construction which optimizes Elmore delay directly, as opposed to heuristically abstracting linear or Elmore delay as in previous approaches. Extensive timing simulations on industry IC and MCM interconnect parameters show that our methods yield trees that significantly improve (by averages of up to 67%) over minimum Steiner routings in terms of delays to identified critical sinks. ERTs also serve as generic high-performance routing trees when no critical sink is specified: for 8-sink nets in standard IC (MCM) technology, we improve average sink delay by 19% (62%) and maximum sink delay by 22% (52%) over the minimum Steiner routing. These approaches provide simple, basic advances over existing performance-driven routing tree constructions. Our results are complemented by a detailed analysis of the accuracy and fidelity of the Elmore delay approximation; we also exactly assess the suboptimality of our heuristic tree constructions. In achieving the latter result, we develop a new characterization of Elmore-optimal routing trees, as well as a decomposition theorem for optimal Steiner trees, which are of independent interest. Kenneth D. Boese, Andrew B. Kahng, Bernard A. McCoy, Gabriel Robins |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1994 | Multi-Way Partitioning Via Spacefilling curves and Dynamic ProgrammingabstractSpectral geometric embeddings of a circuit netlist can lead to fast, high quality m ulti-way partitioning solutions.Furthermore, it has been shown that d-dimensional spectral embeddings (d > 1) are a more powerful tool than single-eigenvector embeddings (d = 1) for multi-way partitioning [2] [4].However, previous methods cannot fully utilize information from the spectral embedding while optimizing netlist-dependent objectives.This work introduces a new multi-way circuit partitioning algorithm called DP-RP.W e begin with a d-dimensional spectral embedding from which a 1-dimensional ordering of the modules is obtained using a space lling curve.The 1dimensional ordering retains useful information from the multi-dimensional embedding while allowing application of ecient algorithms.W e show that for a new Restricted Partitioning formulation, dynamic programming eciently nds optimal solutions in terms of Scaled Cost [4] and can transparently handle userspeci ed cluster size constraints.For 2-way ratio cut partitioning, DP-RP yields an average of 45% improvement o v er KP [4] and EIG1 [6] and 48% improvement o v er KC [ 2 ]. Charles J. Alpert, Andrew B. Kahng |
DAC | 2 |
| 1994 | Rectilinear Steiner Trees with Minimum Elmore DelayabstractWe povide a new theoretical framework for constructing Steiner routing trees with minimum Elmore delay. Earlier work [3, 13] has established Elmore delay as a high fidelity estimate of "physical", i.e., SPICE-computed, signal delay. Previously, however, it was not known how to construct an Elmore delay-optimal Steiner tree. Our main theoretical result is a generalization of Hanan's theorem [11] which limited the number of possible locations of Steiner nodes in an optimal delay rectilinear Steiner tree. Another theoretical result establishes a new decomposition theorem for constructing optimal-delay Steiner trees. We develop a branch-and-bound method, called BB-SORT-C, which exactly minimizes any linear combination of Elmore sink delays; BB-SORT-C is practical for routing small nets and for delimiting the space of achievable routing solutions with respect to Elmore delay. Kenneth D. Boese, Andrew B. Kahng, Bernard A. McCoy, Gabriel Robins |
DAC | 2 |
| 1994 | Delay Analysis of VLSI Interconnections Using the Diffusion Equation ModelabstractThe traditional analysis of signal delay in a transmission line begins with a lossless LC representation, which yields a wave equation governing the system response; 2-port parameters are typically derived and the solution is obtained in the transform domain. In this paper, we begin with a distributed RC line model of the interconnect and analytically solve the resulting diffusion equation for the voltage response. A new closed form expression for voltage response is obtained by incorporating appropriate boundary conditions for interconnect delay analysis. Calculations of 50% and 90% delay times for various cases of interest (e.g., open-ended RC line) give substantially different estimates from those commonly cited in the literature, thus suggesting revised delay estimation methodologies and intuitions for the design of VLSI interconnects. The discussion furthermore provides a unifying treatment of the past three decades of RC interconnect delay analyses. 1 Overview Delay analysis o... Andrew B. Kahng, Sudhakar Muddu |
DAC | 1 |
| 1994 | A general framework for vertex orderings, with applications to netlist clustering
Charles J. Alpert, Andrew B. Kahng |
ICCAD | 2 |
| 1994 | Low-cost single-layer clock trees with exact zero Elmore delay skew
Andrew B. Kahng, Chung-Wen Albert Tsao |
ICCAD | 1 |
| 1994 | On the intrinsic Rent parameter and spectra-based partitioning methodologiesabstractThe complexity of circuit designs has necessitated a top-down approach to layout synthesis. A large body of work shows that a good layout hierarchy, or partitioning tree, as measured by the associated Rent parameter, will correspond to an area-efficient layout. We define the intrinsic Rent parameter of a netlist to be the minimum possible Rent parameter of any partitioning tree for the netlist. Experimental results show that spectra-based ratio cut partitioning algorithms yield partitioning trees with the lowest observed Rent parameter over all benchmarks and over all algorithms tested. For examples where the intrinsic Rent parameter is known, spectral ratio cut partitioning yields a partitioning tree with Rent parameter essentially identical to this theoretical optimum. These results have deep implications with respect to both the choice of partitioning algorithms for top-down layout, as well as new approaches to layout area estimation. The paper concludes with directions for future research, including several promising techniques for fast estimation of the (intrinsic) Rent parameter.> Lars W. Hagen, Andrew B. Kahng, Fadi J. Kurdahi, Champaka Ramachandran |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1993 | Geometric Embeddings for Faster and Better Multi-Way Netlist Partitioningabstractalgorithmsfor k-way circuit partitioning Charles J. Alpert, Andrew B. Kahng |
DAC | 2 |
| 1993 | High-Performance Routing Trees With Identified Critical SinksabstractWe present two critical-sink routing tree (CSRT) constructions which exp!oit criticai-path information that becomes availab[e during timing-driven layout.Our CS-Steiner heuristics with "Giobai Slack Removal" modify traditional Stetner constructions and produce routing trees wtth szgntjicanthj lower criticalsink delays compared with existing performance-driven methods.We also propose a new class of Elrnore routing tree (ERT) constructions, which deratively add tree edges to minimize Elmore delay.This direct optimization of Elmore delay yields trees that improve delays to identified crztical sinks by up to 69% over minimum Steiner routtngs.ERTs also improve performance over such recent methods as [1] [6] when no critical stnks are specified. Kenneth D. Boese, Andrew B. Kahng, Gabriel Robins |
DAC | 2 |
| 1993 | Fidelity and Near-Optimality of Elmore-Based Routing ConstructionsabstractWe address the efficient construction of interconnection trees with near-optimal delays. We study the accuracy and fidelity of easily-computed delay models with respect to detailed simulation (e.g., SPICE-computed delays). We show that Elmore delay minimization (W.C. Elmore, 1948) is a high-fidelity interconnect objective for IC interconnect technologies, and propose a greedy low delay tree (LDT) heuristic which for any monotone delay function can efficiently minimize maximum delay. For comparison, we also generate optimal routing trees (ORTs) with respect to Elmore delay, using branch-and-bound search. Experimental results show that the LDT heuristic approximates ORTs very accurately: for nets with up to seven pins, LDT trees have a maximum sink delay within 2.3% of optimum on average. Moreover, compared with minimum spanning tree constructions, the LDT achieves average reductions in delay of up to 35% depending on the net size.> Kenneth D. Boese, Andrew B. Kahng, Bernard A. McCoy, Gabriel Robins |
ICCD | 2 |
| 1993 | Minimum Density Interconneciton Trees
Charles J. Alpert, Jason Cong, Andrew B. Kahng, Gabriel Robins, Majid Sarrafzadeh |
ISCAS | 3 |
| 1993 | A Direct Combination of the Prim and Dijkstra Constructions for Improved Performance-driven Global Routing
Charles J. Alpert, T. C. Hu, Dennis J.-H. Huang, Andrew B. Kahng |
ISCAS | 4 |
| 1993 | Simulated annealing of neural networks: The 'cooling' strategy reconsidered
Kenneth D. Boese, Andrew B. Kahng |
ISCAS | 2 |
| 1993 | Matching-based methods for high-performance clock routingabstractThe authors point out that minimizing clock skew is important in the design of high-performance VLSI systems. A general clock routing scheme that achieves extremely small clock skews while still using a reasonable amount of wirelength is presented. The routing solution is based on the construction of a binary tree using geometric matching. For cell-based designs, the total wirelength of the clock routing tree is on average within a constant factor of the wirelength in an optimal Steiner tree, and in the worst case is bounded by O( square root l/sub 1/l/sub 2/*1 square root n) for n terminals arbitrarily distributed in the l/sub 1/*l/sub 2/ grid. The bottom-up construction readily extends to general cell layouts, where it also achieves essentially zero clock skew within reasonably bounded total wirelength. The algorithms have been tested on numerous random examples and also on layouts of industrial benchmark circuits. The results are very promising: the clock routing yields near-zero average clock skew while using total wirelength competitive with that used by previously known methods.> Jason Cong, Andrew B. Kahng, Gabriel Robins |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1993 | Optimal robust path planning in general environmentsabstractWe address robust path planning for a mobile agent in a general environment by finding minimum cost source-destination paths having prescribed widths. The main result is a new approach that optimally solves the robust path planning problem using an efficient network flow formulation. Our algorithm represents a significant departure from conventional shortest-path or graph search based methods; it not only handles environments with solid polygonal obstacles, but also generalizes to arbitrary cost maps that may arise in modeling incomplete or uncertain knowledge of the environment. Simple extensions allow us to address higher dimensional problem instances and minimum-surface computations; the latter is a result of independent interest. We use an efficient implementation to exhibit optimal path-planning solutions for a variety of test problems. The paper concludes with open issues and directions for future work.> T. C. Hu, Andrew B. Kahng, Gabriel Robins |
IEEE Trans. Robotics Autom. | 2 |
| 1992 | Net Partitions Yield Better Module Partitions
Jason Cong, Lars W. Hagen, Andrew B. Kahng |
DAC | 3 |
| 1992 | A new approach to effective circuit clusteringabstractIt is pointed out that the complexity of next-generation VLSI systems will exceed the capabilities of top-down layout synthesis algorithms, particularly in netlist partitioning and module placement. Bottom-up clustering is needed to condense the netlist so that the problem size becomes tractable to existing optimization methods. Here, the DS quality measure, a general metric for evaluation of clustering algorithms, is established. The DC metric in turn motivates the RW-ST algorithm, a self-tuning clustering method based on random walks in the circuit netlist. RW-ST efficiently captures a globally good circuit clustering. When incorporated within a two-phase iterative Fiduccia-Mattheyses partitioning strategy, the RW-ST clustering method improves bisection width by an average of 17% over previous matching-based methods.> Lars W. Hagen, Andrew B. Kahng |
ICCAD | 2 |
| 1992 | An Improved Graph-Based FPGA Techology Mapping Algorithm For Delay OptimizationabstractA graph-based technology mapping algorithm, called DAG-Map, for delay optimization in lookup-table-based field programmable gate array (FPGA) designs is presented. The algorithm carries out technology mapping and delay optimization on the entire Boolean network, instead of decomposing it into fanout-free trees and mapping each tree separately as in most previous algorithms. As a preprocessing step, a general algorithm that transforms an arbitrary n-input network into a two-input network with only O(1) factor increase in the network depth is introduced. Also presented is a graph-matching-based technique used as a postprocessing step which optimizes the area without increasing the delay. The DAG-Map algorithm was tested on the MCNC logic synthesis benchmarks. Compared with previous algorithms, it reduces both the network depth and the number of lookup-tables.> Jason Cong, Yuzheng Ding, Andrew B. Kahng, Peter Trajmar, Kuang-Chien Chen |
ICCD | 3 |
| 1992 | Provably good performance-driven global routingabstractThe authors propose a provably good performance-driven global routing algorithm for both cell-based and building-block design. The approach is based on a new bounded-radius minimum routing tree formulation. The authors first present several heuristics with good performance, based on an analog of Prim's minimum spanning tree construction. Next, they give an algorithm which simultaneously minimizes both routing cost and the longest interconnection path, so that both are bounded by small constant factors away from optimal. They also show that geometry helps in routing: in the Manhattan plane, the total wire length for Steiner routing improves to 3/2*(1+(1/ epsilon )) times the optimal Steiner tree cost, while in the Euclidean plane, the total cost is further reduced to (2/ square root 3)*(1+(1/ epsilon )) times optimal. The method generalizes to the case where varying wire length bounds are prescribed for different source-sink paths. Extensive simulations confirm that this approach works well.> Jason Cong, Andrew B. Kahng, Gabriel Robins, Majid Sarrafzadeh, Chak-Kuen Wong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1992 | New spectral methods for ratio cut partitioning and clusteringabstractPartitioning of circuit netlists in VLSI design is considered. It is shown that the second smallest eigenvalue of a matrix derived from the netlist gives a provably good approximation of the optimal ratio cut partition cost. It is also demonstrated that fast Lanczos-type methods for the sparse symmetric eigenvalue problem are a robust basis for computing heuristic ratio cuts based on the eigenvector of this second eigenvalue. Effective clustering methods are an immediate by-product of the second eigenvector computation and are very successful on the difficult input classes proposed in the CAD literature. The intersection graph representation of the circuit netlist is considered, as a basis for partitioning, a heuristic based on spectral ratio cut partitioning of the netlist intersection graph is proposed. The partitioning heuristics were tested on industry benchmark suites, and the results were good in terms of both solution quality and runtime. Several types of algorithmic speedups and directions for future work are discussed.> Lars W. Hagen, Andrew B. Kahng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1992 | A new class of iterative Steiner tree heuristics with good performanceabstractA fast approach to the minimum rectilinear Steiner tree (MRST) problem is presented. The method yields results that reduce wire length by up to 2% to 3% over the previous methods, and is the first heuristic which has been shown to have a performance ratio less than 3/2; in fact, the performance ratio is less than or equal to 4/3 on the entire class of instances where the ratio c(MST)/c(MRST) is exactly equal to 3/2. The algorithm has practical asymptotic complexity owing to an elegant implementation which uses methods from computation geometry and which parallelizes readily. A randomized variation of the algorithm, along with a batched variant, has also proved successful.> Andrew B. Kahng, Gabriel Robins |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1992 | On the performance bounds for a class of rectilinear Steiner tree heuristics in arbitrary dimensionabstractA family of examples on which a large class C of minimum spanning tree-based rectilinear Steiner tree heuristics has a performance ratio arbitrarily close to 3/2 times optimal is given. The class C contains many published heuristics whose worst-case performance ratios were previously unknown. Of particular interest is that C contains two heuristics whose worst-case ratios had been conjectured to be bounded away from 3/2, and the construction also points out an incorrect claim of optimality for one of these heuristics. The examples also force the worst possible behavior in a number of heuristics outside C. The construction generalizes to d dimensions, where the heuristics will have performance ratios of at least 2d - 1/d; this improves the previous lower bound on performance ratio in arbitrary dimension.> Andrew B. Kahng, Gabriel Robins |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1991 | High-Performance Clock Routing Based on Recursive Geometric AatchingabstractMinimizing clock skew is a very important problem in the design of high performance VLSI systems. We present a general clock routing scheme that achieves extremely small clock skews, while still using a reasonable amount of wire length. This routing solution is based on the construction of a binary tree using recursive geometric matching. We show that in the average case the total wire length of the perfect path-balanced tree is within a constant factor of the wire length in an optimal Steiner tree, and that in the worst case, is bounded by O(fi when the n leaves are arbitrarily distributed in the unit square. We tested our algorithm on numerous random examples and also on industrial benchmark circuits and obtained very promising results: our clock routing yields near-zero average clock skew while using similar or even shorter total wire length in comparison with the methods of [7]. Andrew B. Kahng, Jason Cong, Gabriel Robins |
DAC | 1 |
| 1991 | Fast Spectral Methods for Ratio Cut Partitioning and ClusteringabstractThe ratio cut partitioning objective function successfully embodies both the traditional min-cut and equipartition goals of partitioning. Fiduccia-Mattheyses style ratio cut heuristics have achieved cost savings averaging over 39% for circuit partitioning and over 50% for hardware simulation applications. The authors show a theoretical correspondence between the optimal ratio cut partition cost and the second smallest eigenvalue of a particular netlist-derived matrix, and present fast Lanczos-based methods for computing heuristic ratio cuts from the eigenvector of this second eigenvalue. Results are better than those of previous methods, e.g. by an average of 17% for the Primary MCNC benchmarks. An efficient clustering method, also based on the second eigenvector, is very successful on the 'difficult' input classes in the CAD (computer-aided design) literature. Extensions and directions for future work are also considered.> Lars W. Hagen, Andrew B. Kahng |
ICCAD | 2 |
| 1991 | Performance-Driven Global Routing for Cell Based ICsabstractAdvances in VLSI technology and the increased complexity of circuit designs cause performance to become an increasingly important constraint for layout. The issue of delay optimization during the global routing phase is addressed. This problem is formulated as the construction of a bounded-radius spanning tree for a given pointset in the plane, and a family of effective heuristics is presented. This approach has very good empirical performance with respect to total wirelength, and can be smoothly tuned between the competing requirements of minimum delay and minimum total netlength, as confirmed by extensive computational results which confirm this. Extensions can be made to the graph and Steiner versions of the problem, and a number of open problems are described.> Jason Cong, Andrew B. Kahng, Gabriel Robins, Majid Sarrafzadeh, Chak-Kuen Wong |
ICCD | 2 |
| 1991 | An Effective Analog Approach to Steiner RoutingabstractConstruction of a minimum rectilinear Steiner tree (MRST) is a fundamental problem in the physical design of VLSI circuits. The problem is NP-complete, and numerous heuristics have been proposed. A novel analog approach is proposed which intuitively shrinks a bubble around the pins of the signal net until a Steiner tree topology is induced. The method easily maps to parallel neural-style architectures, as well as to fairly generic two-dimensional processor arrays. Extensive simulation results show better performance than virtually all existing MRST approaches. The result is a rare instance where an analog heuristic for an NP-complete problem outperforms existing combinatorial methods, both in time complexity and in average-case performance.> Andrew B. Kahng |
ICCD | 1 |
| 1991 | Optimal algorithms for extracting spatial regularity in images
Andrew B. Kahng, Gabriel Robins |
Pattern Recognit. Lett. | 1 |
| 1990 | A New Class of Steiner Trees Heuristics with Good Performance: The Iterated 1-Steiner-ApproachabstractVirtually all previous methods for the rectilinear Steiner tree problem begin with a minimum spanning tree topology and rearrange edges to induce Steiner points. This study presents a more direct approach: the authors iteratively find optimal Steiner points to be added to the layout. The method gives improved average-case performance, and also avoids the worst-case examples of existing approaches. Sophisticated computational geometry techniques allow efficient and practical implementation, and the method is naturally suited to real-world VLSI regimes where, e.g., via costs can be high. Extensive performance results show almost 3% wirelength reduction over the best existing methods. A number of variants and extensions are described.> Andrew B. Kahng, Gabriel Robins |
ICCAD | 1 |
| 1989 | Fast Hypergraph PartitionabstractWe present a new Ο(n2) heuristic for hypergraph min-cut bipartitioning, an important problem in circuit placement. Fastest previous methods for this problem are Ο(n2 log n). Our approach is based on the intersection graph G dual to the input hypergraph. Paths in G are used to construct partial bipartitions which can be completed optimally. The method is provably good and, in particular, obtains optimum results for “difficult” inputs, i.e., hypergraphs with smaller than expected minimum cutsize. Computational results for a wide range of inputs are also discussed. Andrew B. Kahng |
DAC | 1 |