EDBT 2026 Demo / reviewers in the wild / expert
David G. Chinnery
dblp:04/6478
· DBLP profile ↗
25ranked-venue papers
8as first author
7since 2021 · last 2026
0000-0003-2693-439XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 8 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Invited: Improving Runtime Scaling in the EDA Flow for Designs with Millions of GatesabstractDigital circuits have scaled to billions of gates, but synthesis and automated place-and-route (SAPR) tools are practically limited to several millions of gates due to taking weeks of runtime beyond this design size. Server hardware advances, improvement in software code and algorithms, and parallelism, e.g., multi-threading, have provided modest speedups of about 10x in the past 20 years. However, SAPR software is difficult to parallelize with significant overheads for single-threaded synchronization, mutual exclusion (mutex) locks when reading from and writing to database to avoid collisions, to ensure deterministic results, etc. This work discusses why SAPR runtime scaling has been limited, approaches that have been taken to tackle this and why they have been insufficient, looking toward what is possible in the future. Amdahl's law limits speedups achievable with massively parallel techniques like GPU acceleration, which are practically limited to isolated portions of the SAPR flow. Flow restructuring, multi-objective optimization, and partitioning large designs to run in parallel on separate compute servers provide further opportunities for speedup, but suboptimally vs. optimizing the design as a whole. As a detailed example herein, partitioning to speed up swapping between and reordering of scan chains provides up to 3.5x speedup at the cost of up to 10% increased scan chain wire length vs. the fastest results on hard testcases with high runtimes. David G. Chinnery |
ISPD | 1 |
| 2026 | Introduction to the Special Issue on Advances in Physical Design Automation
Stephan Held, Gracieli Posser, Iris Hui-Ru Jiang, David G. Chinnery |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2023 | EDA for Domain Specific Computing: An Introduction for the PanelabstractThis panel explores domain-specific computing from hardware, software, and electronic design automation (EDA) perspectives. Iris Hui-Ru Jiang, David G. Chinnery |
ISPD | 2 |
| 2023 | Task-Based Parallel Programming for Gate SizingabstractPhysical synthesis engines need to embrace all available parallelism to cope with the increasing complexity of modern designs and still offer high quality of results. To achieve this goal, the involved algorithms need to be expressed in a way that facilitates fast execution time across a range of computing platforms. In this work, we introduce a task-based parallel programming template that can be used for speeding up timing and power optimization. This approach utilizes all available parallelism and enables significant speedup relative to custom multithreaded approaches. Task-based parallelism is applied to all parts of the optimization engine covering also parts that are traditionally executed serially for preserving maximum timing accuracy. Using Taskflow as the parallel programming and execution engine, we achieved a speedup of$1.7\times $to$2.8\times $for gate sizing optimizations on the ISPD13 benchmarks with marginal extra leakage power relative to state-of-the-art multithreaded gate sizers. This result was supported by two dynamic heuristics that restrict the number of examined gate sizes and simplify local timing updates. Both heuristics tradeoff additional runtime reduction with marginal leakage power increases. Dimitrios Mangiras, David G. Chinnery, Giorgos Dimitrakopoulos |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Introduction to the Special Section on Advances in Physical Design Automationabstractintroduction Share on Introduction to the Special Section on Advances in Physical Design Automation Authors: Iris Hui-Ru Jiang National Taiwan University National Taiwan University 0000-0002-4554-3442View Profile , David Chinnery Siemens Digital Industries Software Siemens Digital Industries Software 0000-0003-2693-439XView Profile , Gracieli Posser Cadence Design Systems Cadence Design Systems 0000-0003-4683-3676View Profile , Jens Lienig Dresden University of Technology Dresden University of Technology 0000-0002-2140-4587View Profile Authors Info & Claims ACM Transactions on Design Automation of Electronic SystemsVolume 28Issue 5Article No.: 68pp 1–3https://doi.org/10.1145/3604593Published:09 September 2023Publication History 0citation39DownloadsMetricsTotal Citations0Total Downloads39Last 12 Months39Last 6 weeks39 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Iris Hui-Ru Jiang, David G. Chinnery, Gracieli Posser, Jens Lienig |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2022 | Integrating LR Gate Sizing in an Industrial Place-and-Route FlowabstractLagrangian relaxation (LR) based gate sizing is the state-of-the-art gate-sizing approach. Integrating it within a place-and-route (P&R) tool is difficult as LR needs multiple iterations to converge, requiring very fast timing analysis. Gate-sizing is invoked in many P&R flow steps, so it is also unclear where best to use LR sizing. We detail development of a LR gate sizer for an industrial P&R flow. Software architecture and P&R flow needs are discussed. We summarize how we sped up the LR sizer by 3x to resize a million gates per hour, and ensure multi-threaded results are deterministic. LR sizing experiments at the fast WNS/TNS optimization steps in the flow stages before and after clock tree synthesis (CTS) show excellent results: 10% to 20% setup timing total negative slack (TNS) reduction with 11% to 14% less leakage power, or 1% to 3% lower total power (dynamic power + leakage) with a total power objective, and 1% to 3% lower cell area. Worst negative slack (WNS) also improved in 2/3 of designs in pre-CTS. In the full flow, 5% lower leakage, 1% lower total power, and 0.6% lower cell area can be achieved, with roughly neutral impact on other metrics, compared to a high-effort low-power P&R flow baseline. David G. Chinnery, Ankur Sharma 0001 |
ISPD | 1 |
| 2021 | Autonomous Application of Netlist Transformations Inside Lagrangian Relaxation-Based OptimizationabstractTiming closure is a complex process that involves many iterative optimization steps applied in various phases of the physical design flow. Lagrangian relaxation (LR)-based optimization has been established as a viable approach for this. We extend LR-based optimization by interleaving in each iteration various techniques, such as gate and flip-flop sizing, buffering to fix late and early timing violations, pin swapping, gate merge/split transformations, and useful clock skew. In all cases, locally optimal decisions are made using LR-based cost functions. In each iteration of LR-based optimization, we leverage the multiarmed bandit (MAB) model to automatically pick which optimization heuristic should be applied to the design. The goal is to improve the performance metrics based on the rewards learned from the previous applications of each heuristic and the runtime cost paid for the received reward. The fine-grained combination of an LR-based optimization flow with a statistical recommendation system allows for the autonomous execution of the optimization flow and results in significant quality-of-results improvement relative to the state-of-the-art. More specifically, our flow achieves 17% lower clock period, while also saving 15% power and 6% area, on average, on the TAU2019 benchmarks, as compared to the TAU2019 contest winner, and 25% better leakage power on the ISPD13 benchmarks, as compared to the best reported results. Apostolos Stefanidis, Dimitrios Mangiras, Chrysostomos Nicopoulos, David G. Chinnery, Giorgos Dimitrakopoulos |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Design Optimization by Fine-grained Interleaving of Local Netlist Transformations in Lagrangian RelaxationabstractDesign optimization modifies a netlist with the goal of satisfying the timing constraints at the minimum area and leakage power, without violating any slew or load capacitance constraints. Lagrangian relaxation (LR) based optimization has been established as a viable approach for this. We extend LR-based optimization by interleaving in each iteration techniques such as: gate and flip-flop sizing; buffering to fix late and early timing violations; pin swapping; and useful clock skew. Locally optimal decisions are made using LR-based cost functions, without the need for incremental timing updates. Sub-steps are applied in a balanced manner, accounting for the expected savings and any conflicting timing violations, maximizing the final quality of results under multiple process/operating corners with a reasonable runtime. Experimental results show that our approach achieves better timing, and both lower area and leakage power than the winner of the TAU 2019 contest, on those benchmarks. Apostolos Stefanidis, Dimitrios Mangiras, Chrysostomos Nicopoulos, David G. Chinnery, Giorgos Dimitrakopoulos |
ISPD | 4 |
| 2020 | Fast Lagrangian Relaxation-Based Multithreaded Gate Sizing Using Simple Timing CalibrationsabstractAccurate delay analysis with distributed RC delay can be computationally expensive, and can contribute the majority of the total runtime for gate sizers. Recent works have shown that Lagrangian relaxation (LR)-based gate sizers have produced designs with the lowest power on average. But they are also very slow due to a large number of expensive timing updates spread across several tens of iterations. In this paper, we develop an LR-based discrete gate sizer for fast timing and power reduction. Our gate sizer is multithreaded and is equipped with parallelization enabling techniques, namely mutual exclusion edge (MEE) assignment and directed acyclic graph (DAG)-based netlist traversal (DNT). MEEs are dummy edges assigned to improve load sharing among different threads. DNT facilitates simultaneous resizing of gates belonging to different topological levels. Our Lagrange multiplier update strategy enables rapid convergence of our timing and power recovery algorithms. To reduce the runtime of timing updates, we propose a simple and fast-to-compute effective capacitance model. We further propose mechanisms to calibrate timing models to improve their accuracy. By calibrating the internal timing models only twice, our proposed gate sizing flow facilitates extremely fast design optimization. We benchmark our gate sizer using the ISPD 2012 and 2013 gate sizing contest benchmark suites. Compared to the state-of-the-art gate sizer, our proposed gate sizer is on average 15× faster and the optimized designs have 2.5% higher leakage power. Since we tradeoff timing accuracy for larger runtime speedup, our optimized designs have small timing violations. Ankur Sharma 0001, David G. Chinnery, Tiago Reimann, Sarvesh Bhardwaj, Chris C. N. Chu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Session details: Detailed Routing Contest Results
David G. Chinnery |
ISPD | 1 |
| 2019 | Lagrangian Relaxation Based Gate Sizing With Clock Skew Scheduling - A Fast and Effective ApproachabstractRecent work has established Lagrangian relaxation (LR) based gate sizing as state-of-the-art providing the best power reduction with low run time. Gate sizing has limited potential to reduce the power when the timing constraints are tight. By adjusting the arrival times of clock signals (clock skew scheduling), the timing constraints can be relaxed facilitating more power reduction. Ankur Sharma 0001, David G. Chinnery, Chris C. N. Chu |
ISPD | 2 |
| 2019 | Timing-Driven and Placement-Aware Multibit Register CompositionabstractMultibit register (MBR) composition is an effective and proven method for clock tree power reduction. The proposed MBR composition follows a balanced restructuring approach that is applied after global or detailed placement. Its goal is to minimize the total number of registers in a design, and simplify subsequent clock tree synthesis, while taking care that any potential degradations in timing slack, wire length, or routing congestion do not offset the power benefits of a lighter clock tree. The proposed methodology identifies nearby compatible registers that can be merged without degrading timing, and without reducing the “useful clock skew” potential. These registers are merged, provided that the MBR placement can be legalized according to the proposed simplified physical constraints. A new integer linear programming formulation minimizes the total number of registers in the design. Additional optimization steps give significant reductions in register count and clock tree capacitance, as shown by experimental results on industrial benchmarks that are already rich in MBRs after logic synthesis. These steps include: MBR decomposition; initial allowance of incomplete MBRs, and the partial recovery of them by the end of the flow; and MBR-specific register sizing. Ioannis Seitanidis, Giorgos Dimitrakopoulos, Pavlos M. Mattheakis, Laurent Masse-Navette, David G. Chinnery |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2017 | Timing Driven Incremental Multi-Bit Register Composition Using a Placement-Aware ILP formulationabstractTo reduce clock power, we present a novel timing-driven incremental multi-bit register (MBR) composition methodology for designs that may be rich in MBRs after logic synthesis. It identifies nearby compatible registers that can be merged without degrading timing, and without reducing the "useful clock skew" potential. These registers are merged providing the MBR placement can be legalized according to the proposed simplified physical constraints. A new integer linear programming (ILP) formulation minimizes the total number of registers in the design. It significantly reduces register count and clock capacitance, without adding any timing/routing/placement violations and without increasing the total wire-length of the designs, as shown by experimental results on industrial benchmarks. Ioannis Seitanidis, Giorgos Dimitrakopoulos, Pavlos M. Mattheakis, Laurent Masse-Navette, David G. Chinnery |
DAC | 5 |
| 2017 | Rapid gate sizing with fewer iterations of Lagrangian RelaxationabstractExisting Lagrangian Relaxation (LR) based gate sizers take many iterations to converge to a competitive solution. In this paper, we propose a novel LR based gate sizer which dramatically reduces the number of iterations while achieving a similar reduction in leakage power and meeting the timing constraints. The decrease in the iteration count is enabled by an elegant Lagrange multiplier update strategy for rapid coarse-grained optimization as well as finer-grained timing and power recovery techniques, which allow the coarse-grained optimization to terminate early without compromising the solution quality. Since LR iterations dominate the total runtime, our gate sizer achieves an average speedup of 2.5x in runtime and saves 1% more power compared to the previous fastest work. Ankur Sharma 0001, David G. Chinnery, Shrirang Dhamdhere, Chris C. N. Chu |
ICCAD | 2 |
| 2015 | Fast Lagrangian Relaxation Based Gate Sizing using Multi-ThreadingabstractWe propose techniques to achieve very fast multi-threaded gate-sizing and threshold-voltage swap for leakage power minimization. We focus on multi-threading Lagrangian Relaxation (LR) based gate sizing which has shown both better power savings and better runtime compared to other gate sizing approaches. Our techniques, mutual exclusion edge assignment and directed graph-based netlist traversal, maximize thread execution efficiency to take full advantage of the inherent parallelism when solving the LR subproblem, without compromising the leakage power savings. With 8 threads, our multi-threading techniques achieve on average 5.23x speedup versus our single-threaded (sequential) implementation. This compares well to the maximum achievable speedup of 5.93x by Amdahl's law due to 5% of the execution not being parallelizable. To highlight the problems with load imbalance and poor scheduling, we also propose a simpler approach based on clustering and topological levelby-level netlist traversal, which can achieve only 3.55x speedup. We also propose three simple yet effective enhancements - fast optimal local resizing, early exit policy, and fast greedy timing recovery - to speed up single-threaded LR-based gate-sizing without degrading the leakage power. We test our gate sizer using the ISPD 2012 gate sizing contest benchmarks and guidelines. Compared to other researchers' state-of-the-art LR-based gate sizer, our approach is 1.03x (with 1-thread) and 5.40x (with 8-threads) faster and only 2.2% worse in leakage power. Ankur Sharma 0001, David G. Chinnery, Sarvesh Bhardwaj, Chris C. N. Chu |
ICCAD | 2 |
| 2015 | ISPD 2015 Benchmarks with Fence Regions and Routing Blockages for Detailed-Routing-Driven PlacementabstractThe ISPD~2015 placement-contest benchmarks include all the detailed pin, cell, and wire geometry constraints from the 2014 release, plus Ismail Bustany, David G. Chinnery, Joseph R. Shinnerl, Vladimir Yutsis |
ISPD | 2 |
| 2014 | ISPD 2014 benchmarks with sub-45nm technology rules for detailed-routing-driven placementabstractThe public release of realistic industrial placement benchmarks by IBM and Intel Corporations from 1998--2013 has been crucial to the progress in physical-design algorithms during those years. Direct comparisons of academic tools on these test cases, including widely publicized contests, have spurred researchers to discover faster, more scalable algorithms with significantly improved quality of results. Vladimir Yutsis, Ismail Bustany, David G. Chinnery, Joseph R. Shinnerl, Wen-Hao Liu 0001 |
ISPD | 3 |
| 2013 | High performance and low power design techniques for ASIC and custom in nanometer technologiesabstractTraditionally, synthesized application-specific integrated circuits (ASICs) have been slower and higher power than custom integrated circuits due to a variety of factors. This paper details how this gap has decreased in the past few years. ASICs have adopted higher performance and lower power design techniques with the aid of better CAD tool support. To improve productivity, many full custom designs have migrated to a semi-custom design methodology that is more amenable to the use of standard CAD tools and makes greater use of synthesis. David G. Chinnery |
ISPD | 1 |
| 2005 | Closing the power gap between ASIC and custom: an ASIC perspectiveabstractWe investigate differences in power between application-specific integrated circuits (ASICs) and custom integrated circuits, with examples from 0.6um to 0.13um CMOS. A variety of factors cause synthesizable designs to consume '3 to '7 more power. We discuss the shortcomings of typical synthesis flows, and changes to tools and standard cell libraries needed to reduce power. Using these methods, we believe that the power gap between ASICs and custom circuits can be closed to within 2'. David G. Chinnery, Kurt Keutzer |
DAC | 1 |
| 2005 | Linear programming for sizing, Vth and Vdd assignmentabstractMost circuit sizing tools calculate the tradeoff between each gate's delay and power or area, and then greedily change the gate with the best tradeoff. We show this is suboptimal. Instead we use a linear program to minimize circuit power. The linear program provides a fast and simultaneous analysis of how each gate affects gates it has a path to. Our approach reduces power by up to 30% compared to commercial software, with a 0.13um library. The runtime for posing and solving the linear program scales linearly with circuit size David G. Chinnery, Kurt Keutzer |
ISLPED | 1 |
| 2003 | Low Power Multiplication Algorithm for Switching Activity Reduction through Operand DecompositionabstractA novel low power multiplication algorithm for reducing switching activity through operand decomposition is proposed. Our experimental results show 12% to 18% reduction in logic transitions in both array multipliers and tree multipliers of 32 bits and 64 bits. Similar results are obtained for dynamic power dissipation after logic synthesis. One additional logic gate is required on the critical path for operand decomposition, which corresponds to only an additional 2% to 6% of total delay in these four cases. Thus, the proposed algorithm can be applied to many digital systems where power consumption is a major design constraint. Masayuki Ito, David G. Chinnery, Kurt Keutzer |
ICCD | 2 |
| 2003 | Minimization of dynamic and static power through joint assignment of threshold voltages and sizing optimizationabstractWe describe an optimization strategy for minimizing total power consumption using dual threshold voltage (Vth) technology. Significant power savings are possible by simultaneous assignment of Vth with gate sizing. We propose an efficient algorithm based on linear programming that jointly performs Vth assignment and gate sizing to minimize total power under delay constraints. First, linear programming assigns the optimal amounts of slack to gates based on power-delay sensitivity. Then, an optimal gate configuration, in terms of Vth and transistor sizes, is selected by an exhaustive local search. Benchmark results for the algorithm show 32% reduction in power consumption on average, compared to sizing only power minimization. There is up to a 57% reduction for some circuits. The flow can be extended to dual supply voltage libraries to yield further power savings. Abhijit Davare, Michael Orshansky, David G. Chinnery, Brandon Thompson, Kurt Keutzer |
ISLPED | 4 |
| 2001 | Achieving 550Mhz in an ASIC MethodologyabstractTypically, good automated ASIC designs may be two to five times slower than handcrafted custom designs. At last year's DAC this was examined and causes of the speed gap between custom circuits and ASICs were identified. In particular, faster custom speeds are achieved by a combination of factors: good architecture with well-balanced pipelines; compact logic design; timing overhead minimization; careful floorplanning, partitioning and placement; dynamic logic; post-layout transistor and wire sizing; and speed binning of chips. Closing the speed gap requires improving these same factors in ASICs, as far as possible. In this paper we examine a practical example of how these factors may be improved in ASICs. In particular we show how techniques commonly found in custom design were applied to design a high-speed 550 MHz disk drive read channel in an ASIC design flow. David G. Chinnery, Borivoje Nikolic, Kurt Keutzer |
DAC | 1 |
| 2001 | A Functional Validation Technique: Biased-Random Simulation Guided by Observability-Based CoverageabstractWe present a simulation-based semi-formal verification method for sequential circuits described at the register-transfer level. The method consists of an iterative loop where coverage analysis guides input pattern generation. An observability-based coverage metric is used to identify portions of the circuit not exercised by simulation. A heuristic algorithm then selects probability distributions for biased random input pattern generation that targets non-covered portions. This algorithm is based on an approximate analysis of the circuit modeled as a Markov chain at steady state. Node controllability and observability are estimated using a limited depth reconvergence analysis and an implicit algorithm for manipulating probability distributions and determining steady-state behavior. An optimization algorithm iteratively perturbs the probability distributions of the primary inputs in order to improve estimated coverage. The coverage enhancement achieved by our approach is demonstrated on benchmarks from the ISCAS89 and VIS suites. Serdar Tasiran, Farzan Fallah, David G. Chinnery, Scott J. Weber, Kurt Keutzer |
ICCD | 3 |
| 2000 | Closing the gap between ASIC and custom: an ASIC perspectiveabstractWe investigate the differences in speed between application-specific integrated circuits and custom integrated circuits when each are implemented in the same process technology, with some examples in 0.25 micron CMOS. We first attempt to account for the elements that make the performance different and then examine ways in which tools and methodologies may close the performance gap between application-specific integrated circuits and custom circuits. David G. Chinnery, Kurt Keutzer |
DAC | 1 |