EDBT 2026 Demo / reviewers in the wild / expert
Carl Sechen
dblp:67/1097 · also Carl M. Sechen
· DBLP profile ↗
70ranked-venue papers
4as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 70 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 8 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An SMT-Based Method for Identifying State-Holding Elements in Extracted NetlistsabstractHardware description language (HDL) netlists extracted from reverse-engineered integrated circuits (ICs) are described at the transistor level, thereby obscuring any internal sequential circuitry. Existing methods to extract sequential behavior from transistor-level netlists rely upon prior knowledge of the design, which may not be available. Toward identifying state-holding elements in an extracted netlist without help from any such information, we propose a new methodology which combines graph searching with satisfiability modulo theories (SMT) solving to detect and locate state-holding nets. Aric Fowler, Carl Sechen, Yiorgos Makris |
ITC | 2 |
| 2025 | Modeling Bidirectional Switches for Enabling Logic Equivalence Checking in a Transistor-Level Programmable FabricabstractWe explore the challenges associated with developing a verification solution for a TRAnsistor-level Programmable fabric (TRAP). The TRAP architecture employs bidirectionally operated pass transistors to implement its logic and interconnect network, aiming for high density. However, the existing logic equivalence checking (LEC) methods and tools do not support the primitives necessary to model such transistors in hardware description languages (HDLs). Consequently, verifying the functionality programmed by a given bitstream on TRAP is not inherently feasible. To overcome this limitation, we propose a method that automates the determination of signal flow direction through the bidirectional pass transistors for a given bitstream. Subsequently, we convert the HDL description of the programmed fabric to exclusively utilize unidirectional transistors. This transformation allows us to leverage commercial EDA tools for verifying logic equivalence between the transistor-level HDL representation of the programmed fabric and the post-synthesis gate-level netlist. We have successfully applied the proposed method to verify various benchmark circuits programmed on the TRAP fabric. Apurva Jain, Thomas Broadfoot, Yiorgos Makris, Carl Sechen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | Testing a Transistor-Level Programmable Fabric: Challenges and SolutionsabstractTest vector generation for a TRAnsistor-level Programmable (TRAP) fabric faces a number of feasibility and efficiency challenges. The former are caused by (i) the use of bi-directional pass transistors, which are beyond the capabilities of commercial Automatic Test Pattern Generation (ATPG) tools, and (ii) the design specifics of TRAP, which result in certain stuck-at faults not being logically testable and calling for a quiescent current-based test solution instead. The latter are caused by the fact that ATPG tools are oblivious to (i) the difference between programming bits and regular inputs, which results in lengthy test application times, and (ii) the role that different modules in the architecture of TRAP play in establishing logic circuits, which results in lengthy unguided exploration of a very large functional space to establish appropriate vector justification and response propagation paths. To address these challenges, we explore an array of solutions including (i) employing TRAP instances where bi-directional transistors are replaced by uni-directional ones, (ii) generating custom IDDQ tests, (iii) expressing test application time as the optimization objective of an Integer Linear Program (ILP) formulation, and (iv) leveraging design knowledge, resulting in perfect stuck-at fault coverage of TRAP and an order-of-magnitude savings in test application time. Apurva Jain, Thomas Broadfoot, Carl Sechen, Yiorgos Makris |
VTS | 3 |
| 2023 | Quo Vadis Signal? Automated Directionality Extraction for Post-Programming Verification of a Transistor-Level Programmable FabricabstractWe discuss the challenges related with developing a post-programming verification solution for a TRAnsistor-level Programmable fabric (TRAP). Toward achieving high density, the TRAP architecture employs bidirectionally-operated pass transis-tors in the implementation of its logic and interconnect network. While it is possible to model such transistors through appropriate primitives of hardware description languages (HDL) to enable simulation-based validation, Logic Equivalence Checking (LEC) methods and tools do not support such primitives. As a result, formally verifying the functionality programmed by a given bit-stream on TRAP is not innately possible. To address this limitation, we introduce a method for automatically determining the signal flow direction through bidirectional pass transistors for a given bit-stream and subsequently converting the HDL describing the programmed fabric to consist only of unidirectional transistors. Thereby, commercial EDA tools can be used to check logic equivalence between the transistor-level HDL describing the programmed fabric and the post-synthesis gate-level netlist. Apurva Jain, Thomas Broadfoot, Yiorgos Makris, Carl Sechen |
DATE | 4 |
| 2020 | An Efficient MILP-Based Aging-Aware Floorplanner for Multi-Context Coarse-Grained Runtime Reconfigurable FPGAsabstractShrinking transistor sizes are jeopardizing the reliability of runtime reconfigurable Field Programmable Gate Arrays (FPGAs), making them increasingly sensitive to aging effects such as Negative Bias Temperature Instability (NBTI). This paper introduces a reliability-aware floorplanner which is tailored to multi-context, coarse-grained, runtime reconfigurable architectures (CGRRAs) and seeks to extend their Mean Time to Failure (MTTF) by balancing the usage of processing elements (PEs). The proposed method is based on a Mixed Integer Linear Programming (MILP) formulation, the solution to which produces appropriately-balanced mappings of workload to PEs on the reconfigurable fabric, thereby mitigating aging-induced lifetime degradation. Results demonstrate that, as compared to the default reliability-unaware floorplanning solutions, the proposed method achieves an average MTTF increase of 2.5× without introducing any performance degradation. Mustafa M. Shihab, Yiorgos Makris, Benjamin Carrión Schäfer, Carl Sechen |
DATE | 5 |
| 2020 | Optimal Standard Cell Library Composition for 7nmabstractWe set out to determine the optimal ASIC standard cell library composition for 7nm technology on the basis of power, chip area and delay considerations. We used the publicly available ASAP7, the 7nm FinFET technology from Arizona State University. We also used state-of-the-art commercial computer-aided design (CAD) tools to execute our study. In particular, we synthesized designs using Synopsys' Design Vision and then physical synthesis was performed using Cadence's Innovus. Cell characterization was performed using Synopsys' SiliconSmart. Our results show that standard cell libraries with many fewer cell types than currently used in industry yield the best results. Further, the results suggest that an optimal standard cell library comprises 18 functions along with a rather limited number of drive strengths. Vibhav Kumarswami Salimath, Carl Sechen |
ISCAS | 2 |
| 2020 | CASPER: CAD Framework for a Novel Transistor-Level Programmable FabricabstractA recently proposed TRAnsistor-level Programmable (TRAP) fabric can enable seamless on-die integration of high-density reconfigurable logic with custom ICs. However, state-of-the-art CAD tools are developed for either ASICs or FPGAs and do not support the new architecture. To this end, we present CASPER − a novel CAD framework for implementing designs on the TRAP fabric. CASPER begins with characterizing an ASIC-esque cell library in order to leverage the industry-leading logic synthesis tools for TRAP. We then systematically remodel the TimberWolf and the Versatile Place and Route (VPR) tools to facilitate TRAP-specific design placement and routing, respectively. In addition, we develop a robust programming bitstream generation tool for TRAP. Lastly, we fabricate a 65nm prototype TRAP chip and implement ten ISCAS-85/MCNC benchmark circuits on it. Our evaluation results validate the proposed CAD framework and provide a comparative overhead analysis between TRAP and FPGA. Mustafa M. Shihab, Bharath Ramanidharan, Gaurav Rajavendra Reddy, Jingxiang Tian, William Swartz, Carl Sechen, Yiorgos Makris |
ISCAS | 6 |
| 2020 | ATTEST: Application-Agnostic Testing of a Novel Transistor-Level Programmable FabricabstractA recently introduced TRAnsistor-level Programmable fabric (TRAP) has demonstrated great promise towards seamless unification of high-density reconfigurable logic with Application-Specific Integrated Circuits (ASICs). However, practical deployment of TRAP relies on the development of a comprehensive mechanism for detecting manufacturing defects. Unfortunately, the state-of-the-art test schemes are developed either for ASICs or for Field-Programmable Gate Arrays (FPGAs) and do not support this new transistor-level architecture. To address this limitation, we present a novel application-agnostic test methodology specifically tailored to the TRAP fabric. We first introduce a multi-phase, cascadable scheme to efficiently test the programmable transistors in TRAP’s Logic Elements (LEs). Then, we define the required test patterns for verifying the correct functionality of the built-in D flip-flop, full-adder, and multiplexer of each LE. Next, we present a systematic approach for testing the interconnect network. Lastly, we discuss the limitations in testing the memory cells used for storing the TRAP programming bits and we propose design modifications for improving test coverage. Mustafa M. Shihab, Bharath Ramanidharan, Suraag Sunil Tellakula, Gaurav Rajavendra Reddy, Jingxiang Tian, Carl Sechen, Yiorgos Makris |
VTS | 6 |
| 2019 | Toward an Open-Source Digital Flow: First Learnings from the OpenROAD ProjectabstractWe describe the planned Alpha release of OpenROAD, an open-source end-to-end silicon compiler. OpenROAD will help realize the goal of "democratization of hardware design", by reducing cost, expertise, schedule and risk barriers that confront system designers today. The development of open-source, self-driving design tools is in and of itself a "moon shot" with numerous technical and cultural challenges. The open-source flow incorporates a compatible open-source set of tools that span logic synthesis, floorplanning, placement, clock tree synthesis, global routing and detailed routing. The flow also incorporates analysis and support tools for static timing analysis, parasitic extraction, power integrity analysis, and cloud deployment. We also note several observed challenges, or "lessons learned", with respect to development of open-source EDA tools and flows. Tutu Ajayi, Vidya A. Chhabria, Mateus Fogaça, Soheil Hashemi, Abdelrahman Hosny, Andrew B. Kahng, Jeongsup Lee, Uday Mallappa, Marina Neseem, Geraldo Pradipta, Sherief Reda, Mehdi Saligane, Sachin S. Sapatnekar, Carl Sechen, Mohamed Shalan, William Swartz, Lutong Wang, Zhehong Wang, Mingyu Woo, Bangqi Xu |
DAC | 15 |
| 2019 | Design Obfuscation through Selective Post-Fabrication Transistor-Level ProgrammingabstractWidespread adoption of the fabless business model and utilization of third-party foundries have increased the exposure of sensitive designs to security threats such as intellectual property (IP) theft and integrated circuit (IC) counterfeiting. As a result, concerted interest in various design obfuscation schemes for deterring reverse engineering and/or unauthorized reproduction and usage of ICs has surfaced. To this end, in this paper we present a novel mechanism for structurally obfuscating sensitive parts of a design through post-fabrication TRAnsistor-level Programming (TRAP). We introduce a transistor-level programmable fabric and we discuss its unique advantages towards design obfuscation, as well as a customized CAD framework for seamlessly integrating this fabric in an ASIC design flow. We theoretically analyze the complexity of attacking TRAP-obfuscated designs through both brute-force and intelligent SAT-based attacks and we present a silicon implementation of a platform for experimenting with TRAP. Effectiveness of the proposed method is evaluated through selective obfuscation of various modules of a modern microprocessor design. Results corroborate that, as compared to an FPGA implementation, TRAP-based obfuscation offers superior resistance against both brute-force and oracle-guided SAT attacks, while incurring an order of magnitude less area, power and delay overhead. Mustafa M. Shihab, Jingxiang Tian, Gaurav Rajavendra Reddy, William Swartz, Benjamin Carrión Schäfer, Carl Sechen, Yiorgos Makris |
DATE | 7 |
| 2019 | Functional Obfuscation of Hardware Accelerators through Selective Partial Design Extraction onto an Embedded FPGAabstractThe protection of Intellectual Property (IP) has emerged as one of the most serious areas of concern in the semiconductor industry. To address this issue, we present a method and architecture to map selective portions of a design, given as a behavioral description for High-Level Synthesis (HLS) to a high-security embedded Field-Programmable Gate Array (eFPGA). In this manner, only the end-user has access to the full functionality of the chip. Using six benchmark circuits, we show that our approach is effective. In all cases, the Time-To-Break (TTB) is so long (at least 8 million hours) that for all practical purposes the designs are secure while incurring area overheads of around 5%. Further, latencies were only slightly increased, while the computation times are under one minute. Jingxiang Tian, Mustafa M. Shihab, Gaurav Rajavendra Reddy, William Swartz, Yiorgos Makris, Benjamin Carrión Schäfer, Carl Sechen |
ACM Great Lakes Symposium on VLSI | 8 |
| 2017 | A field programmable transistor array featuring single-cycle partial/full dynamic reconfigurationabstractWe introduce a CMOS computational fabric consisting of carefully arranged regular rows and columns of transistors which can be individually configured and appropriately interconnected in order to implement a target digital circuit. Termed Field Programmable Transistor Array (FPTA), this novel reconfigurable architecture enables several highly-desirable features including (i) simultaneous storage of three configurations along with the ability to dynamically switch between them in a fraction of a single cycle, while retaining the fabric's computational state, (ii) rapid or full modification of a stored configuration in a time proportional to the number of modified configuration bits through the use of hierarchically arranged, high throughput, asynchronously pipelined memory buffers, and (iii) support for libraries containing cells of the same height and variable width, just as in a typical standard cell circuit, thereby simplifying transition from a prototype to a custom IC design. Besides presenting the design details of this fabric in a 130nm technology and demonstrating the aforementioned capabilities, we also briefly discuss the development of a complete CAD flow for programing this fabric and we use numerous benchmark circuits to contrast its area efficiency against a typical FPGA implemented in the same technology node. Jingxiang Tian, Gaurav Rajavendra Reddy, William Swartz, Yiorgos Makris, Carl Sechen |
DATE | 6 |
| 2014 | TonyChopper: a desynchronization packageabstractTonyChopper is an integrated set of tools for digital circuit desynchronization. The core portion of TonyChopper is a tool that reads a gate-level synthesized synchronous digital circuit and transforms it to an asynchronous circuit by implementing a novel desynchronization approach. Pre-layout and post-layout verification tools are also provided in this package. The proposed new asynchronous design method is compatible with conventional synthesis, placement and routing (PnR) and other computer-aided design (CAD) tools. Only a conventional standard cell library is used. Compared to traditional synchronous static CMOS design, the proposed design is highly suitable for very low voltage operation. An auto-sleep strategy is also integrated in the tool for minimizing the leakage power for circuits. Different benchmark circuits were implemented in IBM 130nm technology to show that the design approach used in TonyChopper is highly robust even in the sub-threshold regime. The layout for every benchmark circuit was generated using a Cadence PnR tool. Hspice simulation for both the synchronous benchmark circuit and the desynchronized version provided comparison of the delay, area and leakage power for each benchmark circuit. Monte Carlo simulations were performed for each benchmark circuit to demonstrate high robustness and delay insensitivity for near threshold supply voltages with substantial threshold voltage (VT) variations. Carl Sechen |
ICCAD | 3 |
| 2013 | Low-voltage low-overhead asynchronous logicabstractA new delay-bounded asynchronous logic technique aimed at maximizing reliability at very low voltages is proposed. Compared to previous asynchronous logic approaches, the area and nominal delay overheads are small. Conventional standard cell libraries and conventional logic synthesis tools are used. The bounding delay elements used by the asynchronous controller feature programmable delays that are initially set based on static timing analysis. However, the delay elements are updated on-the-fly during actual operation of the circuit, resulting in strong resiliency even at low voltages and with extreme variations. Several benchmark circuits were implemented with the new asynchronous design flow using the 45nm TI process. Monte Carlo analysis demonstrates the expected resiliency. Compared to the equivalent synchronous circuits, the asynchronous versions have area overheads averaging 40%, although much smaller for large circuits. Nominal delay overheads average about 10%. Akshay Sridharan, Carl Sechen, Roozbeh Jafari |
ISLPED | 2 |
| 2013 | Library-Based Cell-Size Selection Using Extended Logical EffortabstractGiven a synthesized digital integrated circuit comprising interconnected library cells, and assuming arbitrary (continuous) sizes for the cells, experimentally, we have achieved global minimization of the total transistor sizes needed to achieve a delay goal, thus minimizing dynamic power (and reducing leakage power). An accurate table-lookup delay model was developed from the precharacterized industrial standard cell library data by making a formal extension to the concept of logical effort that enables optimization of nMOS and pMOS sizes of a cell separately. To the best of our knowledge, this is the first continuous-cell sizing technique exhibiting optimality based upon a table-lookup delay model. We then developed a new delay-bounded dynamic programming-based algorithm that maps the continuous sizes to the discrete sizes available in the standard cell library, which achieves, for the first time, active area versus delay results close to the continuous results. Parallelism was incorporated into the algorithm to enhance efficiency by leveraging multicore processors. After using state-of-the-art commercial synthesis, the application of our cell-size selection tool results in an active area (the sum of all transistor widths) reduction of 36% (on average) for large contemporary industrial designs. Hiran Tennakoon, Carl Sechen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2012 | Post-synthesis leakage power minimizationabstractWe developed a new post-synthesis algorithm that minimizes leakage power while strictly preserving the delay constraint. A key aspect of the approach is a new threshold voltage (VT) assignment algorithm that employs a cost function that is globally aware of the entire circuit. Thresholds are first raised as much as possible subject to the delay constraint. To further reduce leakage, the delay constraint is then iteratively increased by Δ time units, each time enabling additional cells to have their threshold voltages increased. For each of the iterations, near-optimal cell size selection is applied so as to reacquire the original delay target. The leakage power iteratively reduces to a minimum, and then increases as substantial cell upsizing is required to re-establish the original delay target. We show results for benchmark and commercial circuits using a 40nm cell library in which four threshold voltage options are available. We show that the application of the new leakage power minimization algorithm appreciably reduces leakage power after multi-VTsynthesis by a leading commercial tool, achieving an average post-synthesis leakage reduction of 37% while also reducing total active area and maintaining the original delay target. Carl Sechen |
DATE | 2 |
| 2011 | Power reduction via separate synthesis and physical librariesabstractWe introduce the concept of utilizing two cell libraries, one for synthesis and another for physical design. The physical library consists of only 9 functions, each with several drive and beta ratio options, for a total cell count of 186. We show that synthesis performs better with the inclusion of more complex cells (but only if they are power efficient), we augment the synthesis library to include numerous combinations of the basic 9 functions. The resulting synthesis library consists of a total of 865 cells. Note that these compound cells require only characterization (for a set of drive strengths, but only one beta ratio) and no layout. After design synthesis the compound cells are decomposed back to the basic (9) cells in the physical library. Then cell-size optimization is performed. The entire flow is efficient, with an ability to handle multi-million-gate commercial designs. Applied after state-of-the-art commercial synthesis, the application of a discrete cell-size selection tool, combined with the new dual library approach, results in a typical active area reduction of 40% for large current industrial designs, for the same delay. Ryan Afonso, Hiran Tennakoon, Carl Sechen |
DAC | 4 |
| 2011 | Power reduction via near-optimal library-based cell-size selectionabstractAssuming continuous cell sizes we have robustly achieved global minimization of the total transistor sizes needed to achieve a delay goal, thus minimizing dynamic power (and reducing leakage power). We then developed a feasible branch-and-bound algorithm that maps the continuous sizes to the discrete sizes available in the standard cell library. Results show that a typical library gives results close to the optimal continuous size results. After using state-of-the-art commercial synthesis, the application of our discrete size selection tool results in a dynamic power reduction of 40% (on average) for large industrial designs. Hiran Tennakoon, Carl Sechen |
DATE | 3 |
| 2011 | Power efficient partial product compressionabstractWe present a power efficient structure for partial product compression. Since the full-adder cell is a cornerstone for partial product compression, we determined what are the most power efficient full-adder cells for a wide range of delays. We then developed a new delay-based partial product wiring algorithm that uses the faster XOR-based full adder when one input is a gate delay slower in arriving and otherwise uses the more area efficient mirror-based full adder. We show that this new wiring algorithm is more power efficient than the use of only one type of full adder. Compared to the leading automatic synthesis tool, our approach yields better power efficiencies (lower active area for the same delay) as well as lower delays. For the summation of 16 16-bit vectors, our approach reduces active area by 45% for the same delays produced by the leading synthesis tool. Chiu-wei Pan, Yuanchen Song, Carl Sechen |
ACM Great Lakes Symposium on VLSI | 4 |
| 2008 | Nonconvex Gate Delay Modeling and Delay OptimizationabstractConvex delay models like the Elmore model, the related Logical Effort model, posynomial, and generalized posynomial models have always been favored by researchers, as convexity has a priori guarantees of global optimum solutions. The accuracy of the model may be sacrificed in this quest to generate convex delay models. In this paper, we investigate the use of signomial delay modeling for area/delay optimization. We present a procedure to automatically generate signomial gate delay models by nonlinear least squares fitting. As opposed to posynomial models, signomial models achieve better fits to SPICE generated data. However, signomials are not convex in general. Nevertheless, we show via duality arguments that we obtain near optimum (within 1%) solutions. Our optimization considers beta-ratio constraints, minimum and maximum size constraints for n- and p-transistors, rise/fall delays, and edge rates. The gate sizes for the fastest delay solution for a 44000-cell design, using the IBM 130-nm process, can be achieved in about 16 min of CPU time on a PC, and the area-delay tradeoff curve for 21 points can be generated in about 2 h of CPU time. To the best of our knowledge, this is the first report of using a true signomial delay model and its application to optimum gate sizing. In addition, we give performance details for the automatic data fitting for an 11-function library of static CMOS gates. Hiran Tennakoon, Carl Sechen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | Post-layout comparison of high performance 64b static adders in energy-delay spaceabstractOur objective was to determine the most energy efficient 64 b static CMOS adder architecture, for a range of high-performance delay targets. We examine extensively carry-lookahead (CLA) and carry-select adders with a wide range of tradeoffs in logic levels, fanouts and wiring complexity. We propose sparse CLA adder architectures based on buffering techniques to reduce logic redundancy and improve energy efficiency. All the designs were implemented using an energy-delay layout optimization flow with full RC extraction. Our new 64 b adder designs have a relative delay as low as 9.9 F04 (fanout-offour inverter) delays and promise better scaling for smaller technology nodes. They yield the best energy efficiency for a wide range of delay targets and are 30%, 15% and 7% more energy efficient than full Kogge-Stone, sparse-2 Kogge-Stone and Han-Carlson, respectively, at the fastest points. They consume only about 1/3 the energy of dynamic adders. Carl Sechen |
ICCD | 2 |
| 2006 | Post-layout energy-delay analysis of parallel multipliersabstractThis paper examines parallel multiplier architectures with respect to post-layout energy and delay. Our energy-delay analysis applied to several schemes takes into account the architectural as well as practical circuit implementation issues. Incorporating extracted 3D parasitic loads after layout, our analysis leads to a more realistic conclusion than previous work. A novel parallel-select Booth2 coding method is proposed to improve the performance of the multiplier. Our analysis concludes that the Booth2 method achieves better energy efficiency for large multipliers and for high speed, while the non-Booth method consumes less for small multipliers and longer latencies. Jinyao Zhang, Miodrag Vujkovic, David Wadkins, Carl Sechen |
ISCAS | 4 |
| 2005 | Efficient and accurate gate sizing with piecewise convex delay modelsabstractWe present an efficient and accurate gate sizing tool that employs a novel piecewise convex delay model, handling both rise and fall delays, for static CMOS gates. The delay model is used in a new version of a gate-sizing tool called Forge, which not only exhibits optimality, but also efficiently produces the area versus delay tradeoff curve for a block in one step. Forge includes a realistic delay propagation scheme that combines arrival times and slew-rates. Forge is 6.4X faster than a commercial transistor sizing tool, while achieving better delay targets and uses 28 % less transistor area for specific delay targets, on average. Hiran Tennakoon, Carl Sechen |
DAC | 2 |
| 2004 | Efficient timing closure without timing driven placement and routingabstractWe have developed a design flow from Verilog/VHDL to layout that mitigates the timing closure problem, while requiring no timing driven placement or routing tools. Timing issues are confined to the cell sizer, allowing the placement algorithm to focus solely on net lengths, resulting in superior layout densities and much lower power. The primary enablers to this new technology are: 1) gridded transistor sizing, 2) variable die routing that allows each net to be routed in the shortest possible length, 3) simultaneous cell placement, routing, gate sizing, and clock tree insertion, and 4) an effective incremental (ECO) placement that preserves net lengths from one iteration to the next. The variable die router is the key enabler. Superior placement and routing densities result from this new variable die approach. In addition, each net is routed at or near minimum length, and thus minor placement changes or cell size changes do not materially impact net lengths. Miodrag Vujkovic, David Wadkins, William Swartz, Carl Sechen |
DAC | 4 |
| 2003 | Libraries: lifejacket or straitjacketabstractNo abstract available. Carl Sechen, Barbara Chappel, Jim Hogan, Tadahiko Nakamura, Gregory A. Northrop, Anjaneya Thakar |
DAC | 1 |
| 2003 | Efficient canonical form for Boolean matching of complex functions in large librariesabstractA new algorithm is developed which transforms the truth table or implicant table of a Boolean function into a canonical form under any permutation of inputs. The algorithm is used for Boolean matching for large libraries that contain cells with large numbers of inputs and implicants. The minimum cost canonical form is used as a unique identifier for searching for the cell in the library. The search time is nearly constant if a hash table is used for storing the cells' canonical representations in the library. Experimental results on more than 100000 gates confirm the validity and feasible runtime of the algorithm. Jovanka Ciric, Carl Sechen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2003 | Design and synthesis of dynamic circuitsabstractDynamic circuits are widely used in today's high-performance microprocessors for obtaining timing goals that are not possible using static CMOS circuits. Currently, no commercial tools are able to synthesize dynamic circuits and therefore their design is either completely done by hand or aided by proprietary in-house design tools. This paper describes methodologies and tools for the design and synthesis of dynamic circuits, including general monotonic circuits, which consist of alternating low-skew and high-skew logic gates that may both contain functionality. Synthesis results show standard domino, dynamic-static domino, monotonic static CMOS, zipper CMOS, and footless domino and clock-delayed domino circuits to have average speed improvements of 1.57, 1.66, 1.67, 1.47, 1.71, and 1.60 times over static CMOS, respectively. T. J. Thorp, G. S. Yee, Carl Sechen |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2002 | WTA: waveform-based timing analysis for deep submicron circuitsabstractExisting static timing analyzers make several assumptions about circuits, implicitly trading off accuracy for speed. In this paper we examine the validity of these assumptions, notably the slope approximation to waveforms, single-input transitions, and the choice of a propagating signal based on a single voltage-time point. We provide data on static CMOS gates that show delays obtained in this way can be optimistic by more than 30%. We propose a new approach, Waveform-based Timing Analysis that employs a state-of-the-art circuit simulator as the underlying delay modeler. We show that such an approach can achieve more accurate delays than slope-based timing analyzers at a computation cost that still allows iterations between design modification and delay analysis. Larry McMurchie, Carl Sechen |
ICCAD | 2 |
| 2002 | Gate sizing using Lagrangian relaxation combined with a fast gradient-based pre-processing stepabstractIn this paper, we present Forge, an optimal algorithm for gate sizing using the Elmore delay model. The algorithm utilizes Lagrangian relaxation with a fast gradient-based pre-processing step that provides an effective set of initial Lagrange multipliers. Compared to the previous Lagrangian-based approach, Forge is considerably faster and does not have the inefficiencies due to difficult-to-determine initial conditions and constant factors. We compared the two algorithms on 30 benchmark designs, on a Sun UltraSparc-60 workstation. On average Forge is 200 times faster than the previously published algorithm. We then improved Forge by incorporating a slew-rate-based convex delay model, which handles distinct rise and fall gate delays. We show that Forge is 15 times faster, on average, than the AMPS transistor-sizing tool from Synopsys, while achieving the same delay targets and using similar total transistor area. Hiran Tennakoon, Carl Sechen |
ICCAD | 2 |
| 2002 | Optimized power-delay curve generation for standard cell ICsabstractAn effective way to compare logic techniques, logic families, or cell libraries is by means of power (or area) versus delay plots, since the efficiency of achieving a particular delay is of crucial significance. In this paper we describe a method of producing an optimized power versus delay curve for a combinational circuit. We then describe a method for comparing the relative merits of a set of power versus delay curves for a circuit, each generated with a different cell library. Our results indicate that very few combinational functions need to be in a cell library, at most 11. The power-delay points achieved by Design Compiler from Synopsys using the state-of-the-art Artisan Sage-X library compare unfavorably to our approach. In terms of minimum energy-delay product, our approach is superior by 79% on average. Our approach yields the same delay points with a 107% savings in power consumption, on average. We also show that the specified VDD for a process technology should only be used for the absolute fastest implementations of a circuit. Miodrag Vujkovic, Carl Sechen |
ICCAD | 2 |
| 2002 | Locally clocked pipelines and dynamic logicabstractMicropipelines and most of its variants use a delay-insensitive controller to moderate a pipeline. In search of improved performance, we depart from the delay-insensitive model in favor of a bounded-delay model for the controller. In particular, we demonstrate how a general delay-insensitive controller for level-sensitive pipelines can be improved by assuming a bounded-delay model and taking advantage of delay information to make the controller faster and more efficient. The new control scheme is referred to as locally clocked (LC) control. A highly pipelined logic technique called LC dynamic logic is presented that combines the bounded-delay controller for their comments and suggestions. with a latching dynamic logic gate design. Simulations comparing LC control with its delay-insensitive counterpart are presented. Also, an 8 /spl times/ 8 bit multiplier with a maximum frequency of 715 MHz for a 1 /spl mu/m CMOS process that uses LC dynamic logic is presented. Gregg N. Hoyer, Gin Yee, Carl Sechen |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2001 | Panel: (When) Will FPGAs Kill ASICs?abstractThere was a time - in the dim historical past - when foundries actually made ASICs with only 5000 to 50,000 logic gates. But FPGAs and CPLDs conquered those markets and pushed ASIC silicon toward opportunities with more logic, volume, and speed. Today's largest FPGAs approach the few-million-gate size of a typical ASIC design, and continue to sprout embedded cores, such as CPUs, memories, and interfaces. And given the risks of nonworking nanometer silicon, FPGA costs and time-to-market are looking awfully attractive. So, will FPGAs kill ASICs? ASIC technologists certainly think not. ASICs are themselves sprouting patches of programmable FPGA fabric, and pushing new realms of size and especially speed. New tools claim to have tamed the convergence problems of older ASIC flows. Is the future to be found in a market full of FPGAs with ASIC-like cores? ASICs with FPGA cores? Other exotic hybrids? Our panelists will share their disagreements on these prognostications. Rob A. Rutenbar, Max Baron, Thomas Daniel, Rajeev Jayaraman, Zvi Or-Bach, Jonathan Rose, Carl Sechen |
DAC | 7 |
| 2001 | Automatic datapath tile placement and routingabstractWe report the very first fully automatic datapath tile layout flow. We subdivided the placement process into two steps: a global placement step using simulated annealing, and a new detailed placement step based on extensive modifications we made to the O-tree algorithm. The modifications have enabled the extended O-tree algorithm to handle the rectilinearly shaped transistor chains and gates common in datapath tile layout. We show that datapath tiles can be placed and routed automatically at the transistor level or at the mixed transistor/gate level, achieving results for the very first time that are competitive to those obtained manually by a skilled designer. Tatjana Serdar, Carl Sechen |
DATE | 2 |
| 2001 | Efficient Canonical Form for Boolean Matching of Complex Functions in Large LibrariesabstractA new algorithm is developed which transforms the truth table or implicant table of a Boolean function into a canonical form under any permutation of inputs. The algorithm is used for Boolean matching for large libraries that contain cells with large numbers of inputs and implicants. The minimum cost canonical form is used as a unique identifier for searching for the cell in the library. The search time is nearly constant if a hash table is used for storing the cells' canonical representations in the library. Experimental results on more than 100,000 gates confirm the validity and feasible run-time of the algorithm. Jovanka Ciric, Carl Sechen |
ICCAD | 2 |
| 2001 | Timing- and crosstalk-driven area routingabstractWe present a timing- and crosstalk-driven router for the chip assembly task that is applied between global and detailed routing. Our new approach aims to process the crosstalk and timing constraints by ordering nets and tuning wire spacing in a quantitative way. The new approach fits between global routing and detailed routing along the physical design flow. It is the first to address the timing- and crosstalk-driven area routing problem using crosspoint assignment prior to the detailed routing stage, in contrast to the most previous approaches applied in the post-detailed routing stage. Our new approach enjoys a larger optimization solution space than the previous approaches whose solution space is highly limited by routed geometric constraints. Based on the global routing information, our graph-based optimizer preroutes wires on the global routing grids incrementally. The graph-based optimizer has two stages, net order assignment and space relaxation. A quick capacitance extraction and Elmore delay calculator considering signal switching activities are implemented to find the timing of critical nets and to provide the timing slack database of critical nets. As the graph-based algorithm proceeds, the path delay of critical nets and the timing slack database are updated. During the optimization process, it only optimizes the timing critical paths with negative slack values. The experimental results show a 5%-16% delay reduction for MCNC macrocell benchmark circuits for a 0.25 /spl mu/m process for wire geometric ratio (height/width)=1.0, against a 25% delay reduction if there is infinite space around each metal wire on the same layer. Hsiao-Ping Tseng, Louis K. Scheffer, Carl Sechen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2000 | Delay Minimization and Technology Mapping of Two-Level Structures and Implementation Using Clock-Delayed Domino LogicabstractThis paper presents a new delay minimization and technology mapping algorithm for two-level structures (TLS) implemented using clock-delayed (CD) domino logic. We take advantage of CD domino's high-speed, large fan-in NOR and OR gates to increase the speed of a circuit by partial collapsing. The algorithm is delay-driven and the delays are obtained from a characterized CD domino library. The results on eight combinational MCNC benchmark circuits show an average speed improvement of 89% for CD domino with TLS, compared to static CMOS implementations generated by Synopsys. CD domino with TLS using our tools produced on average 44% faster circuits than CD domino benchmarks minimized and mapped using Synopsys. The delay results for CD domino with TLS were on average 22% better than for standard domino. Jovanka Ciric, Gin Yee, Carl Sechen |
DATE | 3 |
| 2000 | Output Prediction Logic: A High-Performance CMOS Design TechniqueabstractWe present Output Prediction Logic (OPL), a technique that can be applied to conventional CMOS logic families to obtain considerable speedups. When applied to static CMOS, OPL retains the restoring character of the logic family, including its high noise margins. Speedups of 2X to 3X over (optimized) conventional static CMOS are demonstrated for a variety of circuits, ranging from chains of gates, to datapath circuits, ranging from chains of gates, to datapath circuits, and to random logic benchmarks. Such speedups are obtained using identical netlists without remapping. When applied to pseudo-nMOS and dynamic families, in combination with remapping to wide-input NORs, OPL yields speedups of 4X to 5X over static CMOS. Since OPL applied to static CMOS is faster than conventional domino logic, and since it has higher noise margins than domino logic, we believe it will scale much better than domino with future processing technologies. Larry McMurchie, Su Kio, Gin Yee, Tyler Thorp, Carl Sechen |
ICCD | 5 |
| 2000 | Clock-delayed domino for dynamic circuit designabstractClock-delayed (CD) domino is a self-timed dynamic logic family developed to provide single-rail gates with inverting or noninverting outputs. CD domino is a complete logic family and is as easy to design with as static CMOS circuits from a logic design and synthesis perspective. Design tools developed for static CMOS are used as part of a methodology for automating the design of CD domino circuits. The methodology and CD domino's characteristics are demonstrated in the design of a 32-b carry look-ahead adder. The adder was fabricated with MOSIS's 0.8-/spl mu/m CMOS process with scalable CMOS design rules that allow a 1.0-/spl mu/m drawn gate length. Measurements of the adder show a worst case addition of 2.1 ns. The CD domino adder is 1.6/spl times/ faster than a dual-rail domino adder designed with the same cell library and technology. Gin Yee, Carl Sechen |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1999 | AKORD: transistor level and mixed transistor/gate level placement tool for digital data pathsabstractDescribes AKORD, a transistor-level and mixed transistor/gate-level placement tool. AKORD has unique layout capabilities that address the digital data path layout problem. In order to improve communication between the placement and routing steps, new post-placement algorithms were developed: a device re-spacing procedure, an optimization procedure for gate contacts, and a procedure which reduces wire crossovers. AKORD dynamically supports: (1) transistor folding without the usage of device libraries that contain variants of the same device; (2) device merging, including information about optimal transistor chain formation; and (3) well area minimization. Experimental results show that the automated layouts are comparable to skilled manual layouts and that the computation times are quite modest. Tatjana Serdar, Carl Sechen |
ICCAD | 2 |
| 1999 | Design and Synthesis of Monotonic CircuitsabstractWe developed a methodology and tools for synthesizing monotonic networks, which consist of alternating low-skew and high-skew logic gates. By taking advantage of their reduced input capacitance, lower switching thresholds, and efficient implementation for wide complex gates, monotonic circuits can obtain greater performance compared to static CMOS. Our results show standard domino, dynamic-static domino, monotonic static CMOS, and zipper CMOS to have average speed improvements of 1.57, 1.66, 1.67, and 1.47 times over static CMOS, respectively. Tyler Thorp, Gin Yee, Carl Sechen |
ICCD | 3 |
| 1999 | Multilayer pin assignment for macro cell circuitsabstractWe present a multilayer pin (crossing point) assignment algorithm for macro cell circuits. The pin-assignment algorithm takes advantage of a multilayer chip-level global router that we recently developed. Previously reported methods also sought to combine global routing and pin assignment, but their models forced them to use inferior global routing methods. No previous pin-assignment program can handle multilayer layout in which multiple crossing points should be assigned to cell boundaries to fully minimize total routing length and area. In fact, we will show that multiple crossing-point layouts had 27% less total wire length and 9% less area, on average, than single-pin-assignment solutions. Le-Chin Eugene Liu, Carl Sechen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1999 | Multilayer chip-level global routing using an efficient graph-based Steiner tree heuristicabstractWe present a chip-level global router based on a new, more accurate global routing model for the multilayer macro-cell (building block) technology. The routing model uses a three-dimensional mixed directed/undirected routing graph, which provides not only the topological information but also the layer information. The irregular routing graph closely models the multilayer routing problem, so the global router can give an accurate estimate of the routing resources needed. Route-searching is formulated as the Steiner problem in networks (graph Steiner tree problem). Although the Steiner problem in networks is an NP-hard problem, it can generate better routes than other approaches. Previously published Steiner tree heuristics can not handle the complexity of the modern routing graphs. We developed an improved Steiner tree heuristic algorithm which can take advantage of the features of routing graphs. Tested on industrial circuits, our algorithm yields comparable results while having dramatically lower time and space complexities than the leading heuristics. The efficiency and effectiveness of our algorithm make our global router applicable to large industrial circuits, easily handling multilayer problems consisting of 200 macro cells and 10000 nets. While minimizing the wire length, our global router can also minimize the number of vias or solve the routing resource congestion problems. Le-Chin Eugene Liu, Carl Sechen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1999 | A gridless multilayer router for standard cell circuits using CTMcellsabstractWe present a gridless multilayer router suitable for standard cell circuits using central terminal model (CTM) cells. A CTM cell has pins in the middle which split the over-the-cell (OTC) routing region into top and bottom parts. Nets are routed in both the channel (if needed) and OTC by using a channel router. Our router uses a combined constraint graph and tile expansion algorithm. We report the first maze router to take advantage of a graph based algorithm to solve the net ordering problem. Unlike a laggard pure maze router, it has the time efficiency of a fast graph based algorithm and the better results of a maze routing algorithm. This is also the first report of the use of a tile expansion maze router for the variable height routing problem. By variable height routing we mean that a single execution of the routing algorithm will result in a completely routed solution. The only thing that is not determined a priori is the height or number of tracks needed. Our router achieves channelless solutions for the Primary1 circuit by routing over the cell in three layers. It also generates equal or better results compared to the best of the previous channel routers for all the examples we have tried. In fact, Sockeye is the first router to achieve density solutions for the r1, r3, and r4 examples for two layers and for the r3 example for three layers. Hsiao-Ping Tseng, Carl Sechen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1998 | Timing and Crosstalk Driven Area RoutingabstractWe present a timing and crosstalk driven router for the chip assembly task that is applied between global and detailed routing. Our new approach aims to process the crosstalk and timing constraints by ordering nets and tuning wire spacing in a quantitative way. Our graph-based optimizer preroutes wires on the global routing grids incrementally in two stages - net order assignment and space relaxation. The timing delay of each critical path is calculated taking into account interconnect coupling capacitance. The objective is to reduce the delays of critical nets with negative timing slack values, by tuning net ordering and adding extra wire spacing. It shows a remarkable 8.4-25% delay reduction for MCNC benchmarks for wire geometric ratio=2.0, against a 33% delay reduction if interconnect interference disappear. Hsiao-Ping Tseng, Louis K. Scheffer, Carl Sechen |
DAC | 3 |
| 1998 | Domino logic synthesis using complex static gatesabstractArticle Free Access Share on Domino logic synthesis using complex static gates Authors: Tyler Thorp Department of Eledrid Engineering, University of Washington, Seattle, WA Department of Eledrid Engineering, University of Washington, Seattle, WAView Profile , Gin Yee Department of Eledrid Engineering, University of Washington, Seattle, WA Department of Eledrid Engineering, University of Washington, Seattle, WAView Profile , Carl Sechen Department of Eledrid Engineering, University of Washington, Seattle, WA Department of Eledrid Engineering, University of Washington, Seattle, WAView Profile Authors Info & Claims ICCAD '98: Proceedings of the 1998 IEEE/ACM international conference on Computer-aided designNovember 1998 Pages 242–247https://doi.org/10.1145/288548.288620Published:01 November 1998Publication History 8citation496DownloadsMetricsTotal Citations8Total Downloads496Last 12 Months12Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Tyler Thorp, Gin Yee, Carl Sechen |
ICCAD | 3 |
| 1998 | Chip-level area routingabstractWe present a chip-level area router for modern VLSI technologies. The gridless area router can handle any number of layers, as well as rectilinear blockage areas on any layer. A two-stage divide-and-conquer strategy is applied so that the area router can handle very large chips. The first stage includes an area-minimization loop by using an efficient and accurate multi-layer global router. The global router minimizes the chip area while performing the global routing. According to the global routing results, switchboxes are generated for the whole chip area. Then the switchboxes are sent to the second stage for detailed routing, in which a tile-expansion based switchbox router is used. With multi-level rip-up and re-route techniques, the detailed router is shown to be able to complete many difficult switchboxes. The router was tested on the MCNC building block circuits. Our results show better chip areas than the best previously published results. Le-Chin Eugene Liu, Hsiao-Ping Tseng, Carl Sechen |
ISPD | 3 |
| 1997 | The Future of Custom Cell Generation in Physical SynthesisabstractWe present a subjective review of custom cellgeneration methods in the context of future advances instate-of-the-art digital circuit synthesis. In particular, wedescribe three opportunities for coupling circuitoptimization operations with the library developmentprocess. These operations include electrical optimization,technology mapping, and cell level place and route. Martin Lefebvre 0001, David Marple, Carl Sechen |
DAC | 3 |
| 1997 | A parallel standard cell placement algorithmabstractWe present a loosely coupled parallel algorithm for the placement of standard cell integrated circuits. Our algorithm is a derivative of simulated annealing. The implementation of our algorithm is targeted toward networks of Unix workstations. This is the very first reported parallel algorithm for standard cell placement which yields as good or better placement results than its serial version. In addition, it is the first parallel placement algorithm reported which offers nearly linear speed-up for small numbers of processors, in terms of the number of processors (workstations) used, over the serial version. Despite using the rather slow local area network as the only means of interprocessor communication, the processor utilization is quite high, up to 98% for two processors and 90% for six processors. The new parallel algorithm has yielded the best overall results ever reported for the set of MCNC standard cell benchmark circuits. Wern-Jieh Sun, Carl Sechen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1997 | Efficient approximation of symbolic network functions using matroid intersection algorithmsabstractAn efficient and effective approximation strategy is crucial to the success of symbolic analysis of large analog circuits. In this paper we propose a new approximation strategy for the symbolic analysis of linear circuits in the complex frequency domain. The strategy directly generates common spanning trees of a two-graph in decreasing order of tree admittance product, using matroid intersection algorithms. The strategy reduces the total time for computing an approximate symbolic expression in expanded format to polynomial with respect to the circuit size under the assumption that the number of product terms retained in the final expression is polynomial. Experimental results are clearly superior to those reported in previous works. Qicheng Yu, Carl Sechen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1996 | Large Standard Cell Libraries and Their Impact on Layout Area and Circuit PerformancabstractWe present a complete study of layout area and circuit performance as a result of utilizing a large library of standard cells. We built libraries of all possible static CMOS cells having a chain length of up to 7. We refer to a library of all possible cells having a chain length limit of n as sn. Although library s7 has billions of possible cells in it, our technology mapper only selected on the order of 100 of these cells to implement each of the MCNC logic synthesis benchmark circuits. We drew the following conclusions from this study. (1) For three or more layers of metal, using very large libraries (e.g., s7) is optimal in terms of area and delay. (2) For two layers of metal, limiting the library size to s5, but at least s4, is optimal in terms of area and delay. Given that the number of distinct combinational cells in industrial libraries today never exceeds 200 (and usually considerably fewer) and given that even library s4 has 3503 distinct cells, tremendous area savings (without increase in worst case path delay) are readily available by utilizing much larger cell libraries. Specifically, given that library s3 has 87 distinct cells (current industrial libraries typically have no more than this), we surmise that area savings of about 30% can be achieved try using library s7 for three or more metal layers versus any current industrial library. Bingzhong Guan, Carl Sechen |
ICCD | 2 |
| 1996 | Clock-Delayed Domino for Adder and Combinational Logic DesigabstractAn innovative dynamic logic family, clock-delayed (CD) domino, was developed to provide gates with either inverting or non-inverting outputs, and the high speed and layout compactness of dynamic logic. The characteristics of CD domino are demonstrated in two carry lookahead adder designs and three MCNC combinational logic benchmark circuits. The CD domino designs are compared to designs using static CMOS and standard domino logic. A circuit design tool was developed to automate the design of CD domino circuits. Simulations show a 32-bit CD domino adder comprised of four 8-bit full adders to be 30% faster than a 32-bit standard domino adder, anal a 32-bit CD domino adder comprised of a single 32-bit full adder to be 45% faster. In the combinational logic benchmark circuits, complex inverting and non-inverting gates were used to implement C1355, C3540, and b9. The CD domino circuits were 22%, 43% and 34% faster than their static CMOS counterparts of C1355, C3540 and b9, respectively. Gin Yee, Carl Sechen |
ICCD | 2 |
| 1995 | A Method for Finding Good Ashenhurst Decompositions and Its Application to FPGA SynthesisabstractIn this paper, we present an algorithm for finding a good Ashenhurst decomposition of a switching function.Most current methods for performing this type of decomposition are based on the Roth-Karp algorithm.The algorithm presented here is based on finding an optimal cut in a BDD.This algorithm differs from previous decomposition algorithms in that the cut determines the size and composition of the bound set and the free set.Other methods examine all possible bound sets of an arbitrary size.We have applied this method to decomposing functions into sets of -variable functions.This is a required step when implementing a function using a lookup table (LUT) based FPGA.The results compare very favorably to existing implementations of Roth-Karp decomposition methods. Ted Stanion, Carl Sechen |
DAC | 2 |
| 1995 | Timing Driven Placement for Large Standard Cell CircuitsabstractWe present an algorithm for accurately controlling delays during the placement of large standard cell integrated circuits.Previous approaches to timing driven placement could not handle circuits containing 20,000 or more cells and yielded placement qualities which were well short of the state of the art.Our timing optimization algorithm has been added to the placement algorithm which has yielded the best results ever reported on the full set of MCNC benchmark circuits, including a circuit containing more than 100,000 cells.A novel pinpair algorithm controls the delay without the need for user path specification.The timing algorithm is generally applicable to hierarchical, iterative placement methods.Using this algorithm, we present results for the only MCNC standard cell benchmark circuits (fract, struct, and avq.small) for which timing information is available.We decreased the delay of the longest path of circuit fract by 36% at an area cost of only 2.5%.For circuit struct, the delay of the longest path was decreased by 50% at an area cost of 6%.Finally, for the large (22,000 cell) circuit avq.small, the longest path delay was decreased by 28% at an area cost of 6% yet only doubling the execution time.This is the first report of timing driven placement results for any MCNC benchmark circuit. William Swartz, Carl Sechen |
DAC | 2 |
| 1995 | Multiple FPGA Partitioning with Performance OptimizationabstractWe address the problem of partitioning a technology mapped FPGA circuit onto multiple FPGAs of a specific target technology.The physical characteristics of the multiple FPGA system (MFS) pose additional constraints to the circuit partitioning algorithms: the capacity of each FPGA, the timing constraints, the number of I/ Os per FPGA, and the pre-designed interconnection patterns of the MFS.Existing partitioning techniques which minimize just the cut sizes of partitions fail to satisfy the above challenges.We therefore present a rectilinear partitioning algorithm which efficiently and accurately handles timing specifications.The signal path delays are estimated during partitioning using a timing model specific to a multiple FPGA architecture.The model combines all possible delay factors in a system with multiple FPGA chips of a target technology.A new dynamic net-weighting scheme was incorporated to minimize the number of pin-outs for each chip.Finally, we have developed a graph-based global router for pin assignment which can handle the pre-routed connections of our MFS structure.We successfully partitioned the MCNC Xilinx FPGA benchmarks producing 100% routable designs with high utilization levels in all cases.Using the performance optimization capabilities in our approach we have successfully partitioned these benchmarks satisfying the critical path constraints and achieving a significant reduction in the longest path delay.An average reduction of 17% in the longest path delay was achieved at the cost of 5% in total wire length.We have proved the effectiveness of our performance optimization technique by verifying the timing predictions of our partitioner with the actual delays obtained after placement and routing of a partitioned MFS.Partitioning results obtained with the Xilinx mapped MCNC benchmarks are encouraging. Kalapi Roy-Neogi, Carl Sechen |
FPGA | 2 |
| 1995 | Accurate Extraction of Simplified Symbolic Pole/Zero Expressions for Large Analog IC'sabstractWe present the sifting approach and its implementation (AnalogSifter) for symbolic analysis and automatic extraction of symbolic pole and zero expressions for analog integrated circuits. Very compact symbolic expressions for the voltage gains of the two-stage CMOS, 741, and even larger operational amplifiers, were obtained in less than 40 CPU seconds on a SUN SPARCstation 2, and the simplified results match well with the exact numerical values up to the unity gain frequencies in both the magnitude and phase versus frequency plots. The symbolic pole and zero expressions are extracted accurately in the frequency range of interest. All pole and zero expressions may be extracted accurately in individual or clustered symbolic forms by scanning through different frequency ranges. The simplified symbolic pole and zero expressions we obtained in the example presented are within 3% of the exact values. Jer-Jaw Hsu, Carl Sechen |
ISCAS | 2 |
| 1995 | Efficient Approximation of Symbolic Network Function Using Matroid Intersection Algorithms
Qicheng Yu, Carl Sechen |
ISCAS | 2 |
| 1995 | An efficient method for generating exhaustive test setsabstractWe present a new algorithm for generating tests for single stuck line (SSL) faults in combinational logic circuits using a combination of Boolean and path-oriented techniques. We use Boolean techniques employing the ordered binary decision diagram (BDD) to limit the effect of hard faults on the performance of the algorithm. We use path-oriented techniques similar to those in conventional test generation algorithms to limit the number of algebraic operations that must be performed. The test set generated for each fault is exhaustive in the sense that we find all test patterns which make the fault observable at some primary output. We implemented the algorithm as the program TSUNAMI and applied it to the standard ISCAS '85 and ISCAS '89 benchmark circuits. Results indicate that for circuits which are amenable to analysis by BDD's, TSUNAMI performs significantly better than conventional algorithms such as FAN or EST in generating tests for all targeted faults. Moreover, TSUNAMI finds a large set of test vectors for each fault. The program can easily manipulate these sets of vectors to produce test vectors which test many faults. As a result, we are able to compress the sets of vectors into a very small total number of patterns. Ted Stanion, Debashis Bhattacharya, Carl Sechen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1995 | Efficient and effective placement for very large circuitsabstractWe present a new approach to simulated annealing and a new hierarchical algorithm for row-based placement which has obtained the best results ever reported for a large set of MCNC benchmark circuits. Our results indicate that chip area reductions up to 15% are achieved compared with TimberWolfSC v6.0. Our new hierarchical annealing-based placement algorithm (TimberWolfSC v7.0) yields chip area reductions up to 21% while consuming up to 7.5 times less CPU time in comparison to TimberWolfSC v6.0. Furthermore, TimberWolfSC v7.0 produces lower total wire length by an average of 8% than Gordian/Domino, 11% lower wire length than Ritual/Tiger, while using comparable run time. TimberWolfSC v7.0 also supports precise timing driven placement.> Wern-Jieh Sun, Carl Sechen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1994 | Generation of color-constrained spanning trees with application in symbolic circuit analysisabstractProves a property of the second lowest weight spanning tree with color constraint. It enables one to enumerate color-constrained spanning trees in the increasing order of their weights by generalizing existing algorithms for the uncolored problem. A version of the algorithm is implemented in a program for the approximate symbolic analysis of large analog circuits.> Qicheng Yu, Carl Sechen |
Great Lakes Symposium on VLSI | 2 |
| 1994 | A loosely coupled parallel algorithm for standard cell placement
Wern-Jieh Sun, Carl Sechen |
ICCAD | 2 |
| 1994 | Approximate symbolic analysis of large analog integrated circuitsabstractThis paper describes a unified approach to the approximate symbolic analysis of large linearized analog circuits. It combines two new approximation-during-computation strategies with a variation of the classical two-graph tree enumeration method. The first strategy is to generate common trees of the two-graphs, and therefore the product terms in the symbolic network function, in the decreasing order of magnitude. The second approximation strategy is the sensitivity-based simplification of two-graphs, which excludes from the two-graphs many circuit elements that have little effect on the network function being derived. Our approach is therefore able to symbolically analyze much larger analog integrated circuits than previously reported, using complete small signal models for the semiconductor devices. We show accurate yet reasonably sized symbolic network functions for integrated circuits with up to 39 transistors whereas previous approaches were limited to less than 15. Qicheng Yu, Carl Sechen |
ICCAD | 2 |
| 1994 | Boolean division and factorization using binary decision diagramsabstractA method for performing Boolean division and factorization using a new cofactor operation, the interval cofactor, is proposed. This method Is efficiently implemented using BDD's and allows for the use of external and internal don't care sets. As well as generating a normal factored form, the method also generates an extended factored form that allows for the use of the exclusive-OR function in the expression. Using this extended form, much better factorizations may sometimes be found. The method was implemented in Catamount, a logic synthesis system currently under development. The method compares favorably to existing algebraic methods.> Ted Stanion, Carl Sechen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1993 | Maximum projections of don't care conditions in a Boolean networkabstractThe maximum flexibility of a function at a node in a Boolean network is described by the incompletely specified function formed using the union of the satisfiability don't-care set (SDC) and observability don't-care set (ODC) of the node. The normal representation of these sets depends on every variable in the network, can be quite large and may be hard to compute. Usually, we are only interested in the don't care set when restricted to a certain set of variables. We give a formulation for the don't care set of a node in terms of a pre-specified set of variables and prove that this formulation is the maximum projection of the entire don't care set onto the chosen set of variables. This formulation allows computation to start at the target node and proceed by traversing the network outwards. The computation may be stopped at any time yielding a valid subset of the don't care set. Ted Stanion, Carl Sechen |
ICCAD | 2 |
| 1993 | Efficient and effective placement for very large circuitsabstractWe present two major extensions to the implementation of simulated annealing for row-based placement which have enabled it to obtain the best results ever reported for a large set of MCNC benchmark circuits while using the least computation time ever reported for remotely comparable results. Our results indicate that chip area reductions up to 16% can be expected, compared with TimberWolfSC v6.0. Our new hierarchical annealing-based placement program yields total wire length reductions of up to 9% while consuming up to 7.5 times less CPU time in comparison to TimberWolfSC v6.0. In comparison to the Gordian/Domino program, our new program yields total wire lengths which are always lower (up to 9% lower) and our program is always faster for circuits with more than 5000 cells (which represents the range of circuit sizes of interest). Wern-Jieh Sun, Carl Sechen |
ICCAD | 2 |
| 1993 | A new generalized row-based global routerabstractThis paper presents a new generalized row-based global router suitable for standard cell, gate-array, and sea-of-gates integrated circuits. It is the first row-based global router to explicitly minimize chip area. This global router uses adaptive Steiner trees to minimize chip area. The results were vastly improved over typical minimum wire length Steiner trees. This global router automatically adapts to technologies. In addition, optimal feedthrough placement is accomplished using linear assignment. Throughout the algorithm, timing constraints are taken into account. Also, a unique vertical constraint cycle minimization step eases the task for LEA channel routers. Finally, it is shown that this global router outperforms other global routers for all of the MCNC benchmark circuits which were tested. William Swartz, Carl Sechen |
ICCAD | 2 |
| 1990 | New Algorithms for the Placement and Routing of Macro CellsabstractNovel algorithms are described for timing driven placement and routing of rectilinearly shaped macro cells. Algorithms are also presented for the implementation of simulated annealing, based on a theoretically derived statistical annealing schedule. A negative feedback scheme is described that optimizes the relative weighting between the primary objective term and the penalty function terms in the cost function. A placement refinement method has been developed for rectilinear cells which spaces the cells at a density which avoids the need for post-routing compaction. In addition, a detailed routing method has been developed which avoids the classically difficult problem of defining channels for detailed routing. The result for the ami33 benchmark circuit is better than the previously published results.> William Swartz, Carl Sechen |
ICCAD | 2 |
| 1988 | Chip-Planning, Placement, and Global Routing of Macro/Custom Cell Integrated Circuits Using Simulated Annealing
Carl Sechen |
DAC | 1 |
| 1988 | A new global router for row-based layoutabstractA global router for row-based layout styles such as sea-of-gates, gate-array, and standard cell circuits is discussed. It is part of the latest version of TimberWolfSC, a placement and routing package for row-based layout. The algorithm outperformed the UTMC Highland system on two standard benchmark circuits. In tests on ten circuits, the global router produced track counts which were an average of 27% lower than those of the previous TimberWolfSC global router. The router is an average of 30 times faster than the previous algorithm. It has been generalized to handle macro blocks on the chip, equivalent sets of pins, single pins (those without an equivalent), and circuits having many or no built-into-the-cell feeds. Indiscriminate over-the-cell routing is also handled.> Kai-Win Lee, Carl Sechen |
ICCAD | 2 |
| 1988 | An improved objective function for mincut circuit partitioningabstractAn improved objective function has been added to the Kernighan-Lin (1970), Fiduccia-Mattheyses (1982) (KLFM) partitioning algorithm. The time complexity of the enhanced KLFM algorithm remains linear in the number of pins, and there is essentially no change in the CPU time requirements. Based on circuit bipartitioning tests with ten industrial circuits, the number of nets cut was reduced by as much as 55% with the new objective function. The average reduction in nets cut was 38%.> Carl Sechen, Dahe Chen |
ICCAD | 1 |
| 1986 | TimberWolf3.2: a new standard cell placement and global routing packageabstractTimberWolf3.2 is a new standard cell placement and global routing package. The placement and global routing proceed over 3 distinct stages. The general combinatorial optimization technique known as simulated annealing is used during the first two stages of the placement. In the first stage, TimberWolf3.2 places the cells such that the total estimated interconnect cost is minimized. During the second stage, TimberWolf3.2 inserts feed through cells as required and the minimization of the total estimated interconnect cost proceeds again in the manner of simulated annealing. The second stage comes to a close following a global routing step, in which the number of wiring tracks needed is accurately estimated. During the third and final stage, local changes are made to the placement whenever such changes result in a reduction in the number of wiring tracks required. TimberWolf3.2 has achieved area savings ranging from 15 to 75% in experiments on numerous industrial circuits. Carl Sechen, Alberto L. Sangiovanni-Vincentelli |
DAC | 1 |