VLDB 2026 Research / reviewers in the wild / expert
Iris Hui-Ru Jiang
dblp:96/1943
· DBLP profile ↗
109ranked-venue papers
26as first author
32since 2021 · last 2026
0000-0002-4554-3442ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 104 · 26 first-author · 28 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DCPPC: Digital Computation in Programmable Photonic CircuitsabstractPhotonic integrated circuits (PICs) have become a promising alternative to CMOS circuits due to their high speed and energy-saving characteristics. Programmable photonic circuits (PPCs), the field programmable gate array (FPGA) version of PICs, further offers reconfigurability and rapid integration. In this work, we propose a comprehensive methodology for digital computation on PPCs. We first construct the logic building blocks, which serve as unit cells. We devise a novel garbage collection scheme to resolve the signal distortion issue in these cells. A hybrid electronic-photonic scheme and three purely photonic schemes are further proposed to realize the PPC implementation flow of multi-layer optical paths. Experimental validation confirms that the proposed framework successfully synthesizes all 3-input Boolean functions under NPN equivalence.We further extend our approach to support 4-input functions and evaluate its scalability through a case study on implementing a majority voting function, comparing our different proposed schemes. Jun-Wei Liang, Iris Hui-Ru Jiang, Kai-Hsiang Chiu |
ASP-DAC | 2 |
| 2026 | Any-Angle Die-to-Die Routing for Advanced Packages with Asymmetric Pin Row Structures, Via Constraints, and Shielding-Aware ReservationabstractDie-to-die (D2D) routing in advanced packages now faces unprecedented challenges due to extremely dense signal communications, large via size, and strict requirements such as teardrops, staggered vias and full shielding. These constraints severely limit routing resources and necessitate any-angle routing to fully exploit limited space, yet this flexibility drastically increases algorithmic complexity, particularly when dies exhibit irregular, asymmetric pin rows typical in heterogeneous integration. Existing works mostly focus on die-to-substrate (D2S) routing, primarily adopt fixed-angle routing, and most importantly, they all employ sequential route methodologies. As a result, they cannot handle the tight global resource coupling and geometric irregularity of dense D2D scenarios, leading to inferior performance in modern heterogeneous package designs. This paper introduces a global, concurrent any-angle D2D routing framework that unifies routing and via planning, directly incorporates staggered via and teardrop constraints, handles arbitrary asymmetric die structures, and reserves space for full shielding. Experimental results on industrial-inspired D2D benchmarks show that our approach achieves 100% routability in dense regimes, while also accommodating full shielding, reducing total wirelength and maintaining near-linear runtime scalability. Hsin-Tzu Chang, Iris Hui-Ru Jiang, Hua-Yu Chang, Chun-Hao Lai |
ISPD | 2 |
| 2026 | Introduction to the Special Issue on Advances in Physical Design Automation
Stephan Held, Gracieli Posser, Iris Hui-Ru Jiang, David G. Chinnery |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2026 | TriHOT: Triangular and Hexagonal Norm Based Timing-Driven Optical Routing with Wavelength Division MultiplexingabstractAs semiconductor technology continues scaling, interconnect delay becomes a major bottleneck for circuit performance. On-chip optical interconnect with wavelength division multiplexing (WDM) is a promising alternative due to its high speed, broad bandwidth, and energy efficiency. Previous work estimates the timing of an optical interconnect in terms of \(L_1\) or \(L_2\) norm and simplifies its timing gain based on a distance threshold. This estimation and simplification cannot capture the essence of optical routing and the impact of WDM, and may cause timing degradation. In this work, we propose a novel timing model based on a unified triangular and hexagonal norm and establish a wirelength lower bound to measure the routing quality. Moreover, we reduce the WDM-aware optical routing to the Maximum Weighted Independent Set (MWIS) problem and approach it by weighted interval scheduling in linear time-space. Experimental results show that our method outperforms the state-of-the-art in terms of timing, wirelength, and transmission loss with comparable runtime. Jun-Wei Liang, Iris Hui-Ru Jiang |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2026 | Layout Decomposition and Printing Time Optimization for Inkjet-Printed ElectronicsabstractInkjet-printed electronics is a low-cost option for large-scale production. To avoid manufacturing defects, recent research has considered design constraints, such as Laplace and proximity conflicts, decomposed the layouts into different layers, and printed them sequentially. The state-of-the-art work reduced the manufacturing time by optimizing the number of layers and drying time. In this work, we aim to enhance manufacturing efficiency from a new angle, concurrently optimizing the printing time and the layout decomposition of inkjet-printed electronics. We propose an integer linear programming formulation and a dynamic programming algorithm to determine layout decomposition and layer assignment and to estimate the total printing time by carefully considering printing characteristics and design constraints. Experimental results demonstrate significant reductions in overall printing time, leading to improved fabrication efficiency. Meng Lian 0001, Hu Peng, Bernhard Wolfrum, Tsun-Ming Tseng, Iris Hui-Ru Jiang |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2025 | Harrow: Synthesis of Optical Logic Circuits via Harmonic Mean and Integer PartitionabstractWith the advancement of high-speed and energyefficient optical interconnect and computation, photonic integrated circuits (PICs) have become a promising alternative to traditional CMOS circuits. A PIC can be synthesized by mapping the binary decision diagram (BDD) of target functions to optical switches and combiners. However, excessive signal attenuation along the light propagation may require extra optical-electrical signal conversion, thus introducing unwanted delays. In this paper, we aim to overcome this deficiency during logic synthesis: First, we optimize the signal efficiency factor by applying the concept of harmonic means to optimize DC combiners. Second, we eliminate redundant combiners by integer partition. Moreover, we properly arrange these proposed techniques in an optimal sequence of operations to form our main framework. Experimental results show that our framework outperforms the state of the art in terms of efficiency factor. Jun-Wei Liang, Iris Hui-Ru Jiang, Kai-Hsiang Chiu |
DAC | 2 |
| 2025 | Generative Model Based Standard Cell Timing Library CharacterizationabstractAccurate cell timing characterization is essential, on which static timing analysis relies to verify timing performance and ensure design robustness across various PVT conditions (corners). The corner explosion in modern design amplifies the efficiency and scalability challenge for accurate characterization. However, the conventional characterization approach of SPICE simulation alone becomes prohibitively expensive due to the increasing computational complexity and the amount of characterized data. In this paper, we view the characterization problem from a generative modeling perspective to tackle the efficiency and scalability challenge. With a hybrid of generative adversarial network (GAN) and autoencoder, our generative model learns and generalizes among various timing arcs and corners. Experimental results demonstrate that the proposed framework achieves high accuracy and extensibility while reducing the runtime significantly. Hao-Yu Wu, Hsin-Tzu Chang, Shiuan-Yun Ding, Iris Hui-Ru Jiang, Benson Tsao, Vinson Wu, Wei-Kai Shih |
DAC | 4 |
| 2025 | Multi-Stage CSM Timing Waveform Propagation Accelerated by NLDM AssistanceabstractStatic timing analysis (STA) is essential for timing closure. To address the complicated effects emerging at advanced technology nodes, the Current Source Model (CSM) has been developed to compute timing waveforms for timing propagation. Compared with Non-Linear Delay Model (NLDM), CSM provides superior accuracy but suffers from the efficiency and scalability issue. In this paper, we propose a multi-stage CSM timing propagation framework with three acceleration techniques with the assistance of NLDM. Our acceleration techniques are general and compatible with any CSM-based STA engine. Experimental results demonstrate the effectiveness of our acceleration techniques: Compared with CSM-based analysis, we achieve 4× speedups with only 0.4% accuracy loss. Shih-Kai Lee, Pei-Yu Lee, Iris Hui-Ru Jiang |
ISPD | 3 |
| 2025 | Passing-Order Decision for Three-to-Two Lane Merging of Connected and Autonomous Vehicles
Cheng-Pei Chien, Ben-Hau Chia, Ching-Yun Chang, Shang-Chien Lin, Iris Hui-Ru Jiang, Changliu Liu, Chung-Wei Lin |
RTCSA | 5 |
| 2025 | Graceful Register Clustering and Rebanking for Power and Timing BalancingabstractAs dynamic power has become the bottleneck to achieving low power, clock power reduction is crucial in modern IC design. Register clustering can effectively save clock power because of significantly reducing the number of clock sinks and register pin capacitance, clock routed wirelength, and the number of clock buffers. In this article, we propose effective mean shift to naturally form clusters according to register distribution without placement disruption. Effective mean shift fulfills the requirements to be a good register clustering algorithm because it needs no prespecified number of clusters, is insensitive to initializations, is robust to outliers, is tolerant of various register distributions, is efficient and scalable, and balances clock power reduction against timing degradation. Furthermore, we devise a rebanking mechanism to legalize the clustering result according to the adopted multibit register library. Experimental results show that our approach achieves superior power and timing balancing. Tung-Wei Lin, Ya-Chu Chang, Hao-Yu Wu, Iris Hui-Ru Jiang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | Multi-Row Guiding Template Design for Lamellar Directed Self-Assembly with Self-Aligned Via ProcessabstractDirected self-assembly (DSA) of block copolymers can generate tiny and dense layout features, holding great potential for patterning vias and contacts at advanced nodes. Existing studies mainly focused on guiding template design for cylindrical DSA, but by leveraging self-aligned via process, lamellar DSA can form vias to be immune to placement errors and free of a uniform pitch between vias, which cylindrical DSA suffers from. The state-of-the-art guiding template design for lamellar DSA can handle only single-row templates, thus limiting the flexibility of via grouping. Therefore, in this article, we explore further and propose a novel and general multi-row guiding template design approach. Experimental results show that our approach outperforms the state-of-the-art work on both mask conflicts and short guiding templates, and requires much less computation time. Kang-Ting Fan, Iris Hui-Ru Jiang |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2024 | Modern Fixed-Outline Floorplanning with Rectilinear Soft ModulesabstractTo better utilize space and reduce wirelength within a fixed-outline with preplaced modules, modern floorplanning is desired to be able to handle rectilinear soft modules. Nevertheless, the induced special shape constraints have not been fully explored in the literature. In this paper, we propose a novel analytical-based approach to address the challenges of fixed-outline, preplaced modules, and rectilinear soft modules. Unlike most previous work, which abstracts modules as circles during global floorplanning, we treat modules as shape-adjustable rectangles and propose a differentiable shape mechanism to capture the impact of shaping on the floorplan quality. For legalization, we first construct an overlap graph to extract the neighborhood of overlaps, modules, and whitespaces. Then, a shortest-path based algorithm effectively migrates area from overlaps through a chain of modules to whitespaces while carefully considering the shape constraints. Finally, we iteratively expand and shrink the bounding boxes of modules to further refine the wirelength by improving area utilization. Our legalization and refinement allows rectilinear shapes to form naturally. Based on the experiments conducted on the GSRC, MCNC, and CAD contest benchmark suites, our results show that our approach achieves superior wirelength and runtime to the state-of-the-art works and the contest winning team, demonstrating its effectiveness and efficiency. Yuyang Chen 0004, Tzu-Han Hsu, Iris Hui-Ru Jiang, Tung-Chieh Chen, Tai-Chen Chen, Hua-Yu Chang |
ICCAD | 4 |
| 2024 | Slack Redistributed Register Clustering with Mixed-Driving Strength Multi-bit Flip-FlopsabstractRegister clustering is an effective technique for suppressing the increasing dynamic power ratio in modern IC design. By clustering registers (flip-flops) into multi-bit flip-flops (MBFFs), clock circuitry can be shared, and the number of clock sinks and buffers can be lowered, thereby reducing power consumption. Recently, the use of mixed-driving strength MBFFs has provided more flexibility for power and timing optimization. Nevertheless, existing register clustering methods usually employ evenly distributed and invariant path slack strategies. Unlike them, in this work, we propose a register clustering algorithm with slack redistribution at the post-placement stage. Our approach allows registers to borrow slack from connected paths, creates the possibility to cluster with neighboring maximal cliques, and releases extra slack. An adaptive interval graph based on the red-black tree is developed to efficiently adapt timing feasible regions of flip-flops for slack redistribution. An attraction-repulsion force model is tailored to wisely select flip-flops to be included in each MBFF. Experimental results show that our approach outperforms state-of-the-art work in terms of clock power reduction, timing balancing, and runtime. Hao-Yu Wu, Iris Hui-Ru Jiang, Cheng-Hong Tsai, Chien-Cheng Wu |
ISPD | 3 |
| 2024 | Power Sub-Mesh Construction in Multiple Power Domain Design with IR Drop and Routability OptimizationabstractMultiple power domain design is prevalent for achieving aggressive power savings. In such design, power delivery to cross-domain cells poses a tough challenge at advanced technology nodes because of the stringent IR drop constraint and the routing resource competition between the secondary power routing and regular signal routing. Nevertheless, this challenge was rarely mentioned and studied in recent literature. Therefore, in this paper, we explore power sub-mesh construction to mitigate the IR drop issue for cross-domain cells and minimize its routing overhead. With the aid of physical, power, and timing related features, we train one IR drop prediction model and one design rule violation prediction model under power sub-meshes of various densities. The trained models effectively guide sub-mesh construction for cross-domain cells to budget the routing resource usage on secondary power routing and signal routing. Our experiments are conducted on industrial mobile designs manufactured by a 6nm process. Experimental results show that IR drop of cross-domain cells, the routing resource usage, and timing QoR are promising after our proposed methodology is applied. Chien-Pang Lu, Iris Hui-Ru Jiang, Chung-Ching Peng, Mohd Mawardi Mohd Razha, Alessandro Uber |
ISPD | 2 |
| 2024 | Novel Airgap Insertion and Layer Reassignment for Timing Optimization Guided by Slack DependencyabstractBEOL with airgap technology is an alternative metallization option with promising performance, electrical yield and reliability to explore at 2nm node and beyond. Airgaps form cavities in inter-metal dielectrics (IMD) between interconnects. The ultra-low dielectric constant reduces line-to-line capacitance, thus shortening the interconnect delay. The shortened interconnect delay is beneficial to setup timing but harmful to hold timing. To minimize the additional manufacturing cost, the number of metal layers that accommodate airgaps is practically limited. Hence, circuit timing optimization at post routing can be achieved by wisely performing airgap insertion and layer reassignment to timing critical nets. In this paper, we present a novel and fast airgap insertion approach for timing optimization. A Slack Dependency Graph (SDG) is constructed to view the timing slack relationship of a circuit with path segments. With the global view provided by SDG, we can avoid ineffective optimizations. Our Linear Programming (LP) formulation simultaneously solves airgap insertion and layer reassignment and allows a flexible amount of airgap to be inserted. Both SDG update and LP solving can be done extremely fast. Experimental results show that our approach outperforms the state-of-the-art work on both total negative slack (TNS) and worst negative slack (WNS) with more than 89× speedup. Wei-Chen Tai, Min-Hsien Chung, Iris Hui-Ru Jiang |
ISPD | 3 |
| 2024 | Multi-Corner Timing Macro Modeling With Neural Collaborative Filtering From Recommendation Systems PerspectiveabstractTiming macro modeling has been widely employed to enhance the efficiency and accuracy of parallel and hierarchical timing analysis. However, existing studies primarily focused on generating an accurate and compact timing macro model for single-corner libraries, making it difficult to adapt these approaches to multi-corner situations. This either incurs substantial engineering effort or results in significant performance degradation. To tackle this challenge, we offer a fresh perspective on the timing macro modeling problem by drawing inspiration from recommendation systems and formulating it as a matrix completion task. We propose a neural collaborative filtering-based framework capable of capturing the convoluted relationships between circuit pins and timing corners. This framework enables the precise identification of timing variant regions across different corners. Additionally, we design several training features and implement various training techniques to enhance precision. Experimental results show that our framework reduces model sizes by more than 10% compared to state-of-the-art single-corner approaches, while maintaining competitive timing accuracy and exhibiting significant runtime improvements. Furthermore, when applied to unseen corners, our framework consistently delivers superior performance, demonstrating its potential for use in off-corner chiplets in a heterogeneous integration system. Kevin Kai-Chun Chang, Guan-Ting Liu, Chun-Yao Chiang, Pei-Yu Lee, Iris Hui-Ru Jiang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | Graph-Based Deadlock Analysis and Prevention for Robust Intelligent Intersection ManagementabstractIntersection management systems, with the assistance of vehicular networks and autonomous vehicles, have the potential to perform traffic control more precisely than contemporary signalized intersections. However, as infrastructural intersection management controllers do not directly activate motions of vehicles, it is possible that the vehicles fail to follow the instructions from controllers, undermining system properties such as deadlock-freeness and traffic performance. In this article, we consider a class of robustness issues, the time violations, which stem from possible discrepancies between scheduled orders and real executions. We refine a graph-based intersection model to build our theoretical foundations and analyze potential deadlocks and their resolvability. We develop solutions that mitigate negative effects of time violations. In particular, we propose a Robustness-Aware Greedy Scheduling algorithm for robust scheduling and evaluate the deadlock-free robustness of different intersection models and scheduling algorithms. Experimental results show that the Robustness-Aware Greedy Scheduling algorithm is able to significantly improve robustness and keep a good balance with traffic performance. Kai-En Lin, Kuan-Chun Wang, Yu-Heng Chen, Li-Heng Lin, Ying-Hua Lee, Chung-Wei Lin, Iris Hui-Ru Jiang |
ACM Trans. Cyber Phys. Syst. | 7 |
| 2023 | Lightning Talk: All Routes to Timing ClosureabstractTiming analysis and optimization is essential for each stage throughout the entire design flow to achieve timing closure. This talk provides a retrospective and prospective study to highlight two categories of emerging timing challenges: 1) analysis efficiency and scalability, and 2) advanced process effect on timing. We survey recent advances for handling these challenges and provide future research directions in timing. Iris Hui-Ru Jiang |
DAC | 1 |
| 2023 | EDA for Domain Specific Computing: An Introduction for the PanelabstractThis panel explores domain-specific computing from hardware, software, and electronic design automation (EDA) perspectives. Iris Hui-Ru Jiang, David G. Chinnery |
ISPD | 1 |
| 2023 | Introduction to the Special Section on Advances in Physical Design Automationabstractintroduction Share on Introduction to the Special Section on Advances in Physical Design Automation Authors: Iris Hui-Ru Jiang National Taiwan University National Taiwan University 0000-0002-4554-3442View Profile , David Chinnery Siemens Digital Industries Software Siemens Digital Industries Software 0000-0003-2693-439XView Profile , Gracieli Posser Cadence Design Systems Cadence Design Systems 0000-0003-4683-3676View Profile , Jens Lienig Dresden University of Technology Dresden University of Technology 0000-0002-2140-4587View Profile Authors Info & Claims ACM Transactions on Design Automation of Electronic SystemsVolume 28Issue 5Article No.: 68pp 1–3https://doi.org/10.1145/3604593Published:09 September 2023Publication History 0citation39DownloadsMetricsTotal Citations0Total Downloads39Last 12 Months39Last 6 weeks39 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Iris Hui-Ru Jiang, David G. Chinnery, Gracieli Posser, Jens Lienig |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2022 | Timing macro modeling with graph neural networksabstractDue to rapidly growing design complexity, timing macro modeling has been widely adopted to enable hierarchical and parallel timing analysis. The main challenge of timing macro modeling is to identify timing variant pins for achieving high timing accuracy while keeping a compact model size. To tackle this challenge, prior work applied ad-hoc techniques and threshold setting. In this work, we present a novel timing macro modeling approach based on graph neural networks (GNNs). A timing sensitivity metric is proposed to precisely evaluate the influence of each pin on the timing accuracy. Based on the timing sensitivity data and the circuit topology, the GNN model can effectively learn and capture timing variant pins. Experimental results show that our GNN-based framework reduces 10% model sizes while preserving the same timing accuracy as the state-of-the-art. Furthermore, taking common path pessimism removal (CPPR) as an example, the generality and applicability of our framework on various timing analysis models and modes are also validated empirically. Kevin Kai-Chun Chang, Chun-Yao Chiang, Pei-Yu Lee, Iris Hui-Ru Jiang |
DAC | 4 |
| 2022 | Many-Layer Hotspot Detection by Layer-Attentioned Visual Question AnsweringabstractExploring hotspot patterns and correcting them as early as possible is crucial to guarantee yield and manufacturability. Existing hotspot detection and pattern classification methods consider only the geometry on one single layer or one main layer with adjacent layers. In this paper, we investigate the linkage between many-layer hotspot patterns and potentially induced defect types. We first cast the many-layer critical hotspot pattern extraction task as a visual question answering (VQA) problem: Considering a many-layer layout pattern an image and a defect type a question, we devise a layer-attentioned VQA model to answer whether the pattern is critical to the queried defect type. Furthermore, our layer attention mechanism attempts to identify the relevance of each layer for different defect types. Experimental results demonstrate that the proposed model has superior question-answering ability for modern layouts with more than thirty layout layers. Yen-Shuo Chen, Iris Hui-Ru Jiang |
DATE | 2 |
| 2022 | Deadlock Analysis and Prevention for Intersection Management Based on Colored Timed Petri NetsabstractWe propose a Colored Timed Petri Net (CTPN) based model for intersection management. With the expressiveness of the CTPN-based model, we can consider timing, vehicle-specific information, and different types of vehicles. We then design deadlock-free policies and guarantee deadlock-freeness for intersection management. To the best of our knowledge, this is the first work on CTPN-based deadlock analysis and prevention for intersection management. Tsung-Lin Tsou, Chung-Wei Lin, Iris Hui-Ru Jiang |
DATE | 3 |
| 2022 | Sub-Resolution Assist Feature Generation with Reinforcement Learning and Transfer LearningabstractAs modern photolithography feature sizes continue to shrink, sub-resolution assist feature (SRAF) generation has become a key resolution enhancement technique to improve the manufacturing process window. State-of-the-art works resort to machine learning to overcome the deficiencies of model-based and rule-based approaches. Nevertheless, these machine learning-based methods do not consider or implicitly consider the optical interference between SRAFs, and highly rely on post-processing to satisfy SRAF mask manufacturing rules. In this paper, we are the first to generate SRAFs using reinforcement learning to address SRAF interference and produce mask-rule-compliant results directly. In this way, our two-phase learning enables us to emulate the style of model-based SRAFs while further improving the process variation (PV) band. A state alignment and action transformation mechanism is proposed to achieve orientation equivariance while expediting the training process. We also propose a transfer learning framework, allowing SRAF generation under different light sources without retraining the model. Compared with state-of-the-art works, our method improves the solution quality in terms of PV band and edge placement error (EPE) while reducing the overall runtime. Guan-Ting Liu, Wei-Chen Tai, Iris Hui-Ru Jiang, James P. Shiely, Pu-Jen Cheng |
ICCAD | 4 |
| 2022 | Intelligent Design Automation for Heterogeneous IntegrationabstractAs the design complexity grows dramatically in modern circuit designs, 2.5D/3D heterogeneous integration (HI) becomes effective for system performance, power, and cost optimization, providing promising solutions to the increasing cost of more-Moore scaling. In this talk, we investigate the chip, package, and board co-design methodology with advanced packages and optical communication considering essential issues on physical design, electrical, thermal, and mechanical effects, timing, and testing, and suggest future research opportunities. Layout: A robust and vertically integrated physical design flow for HI design is needed. We address chip-, package-, and board-level component planning, package-level RDL routing, board-level routing, optical routing, and placement and routing considering warpage and thermal effects. Timing: New chip-level and cross-chip timing analysis techniques are desired. We address timing propagation under current source delay model (CSM), timing analysis and optimization for optical-electrical routing, multi-corner multi-mode analysis for HI, hierarchical MCMM analysis. Testing: The scope covers functional-like test generation, System-in-Package (SiP) online testing, photonic integrated circuits (PIC) testing and design-for-test (DfT), etc. Integration: We shall address chip, package, and board co-design considering multi-domain physics, including physical, electrical, thermal, mechanical, and optical effects and optimization. Iris Hui-Ru Jiang, Yao-Wen Chang, Jiun-Lang Huang, Charlie Chung-Ping Chen |
ISPD | 1 |
| 2022 | Clock Design Methodology for Energy and Computation Efficient Bitcoin Mining MachinesabstractBitcoin mining machines become a new driving force to push the physical limitation of semiconductor process technology. Instead of peak performance, mining machines pursue energy and computation efficiency of implementing cryptographic hash functions. Therefore, the state-of-the-art ASIC design of mining machines adopts near-threshold computing, deep pipelines, and uni-directional data flow. According to these design properties, in this paper, we propose a novel clock reversing tree design methodology for bitcoin mining machines. In the clock reversing tree, the clock of global tree is fed from the last pipeline stage backward to the first one, and the clock latency difference between the local clock roots of two consecutive stages maintains a constant delay. The local tree of each stage is well balanced and keeps the same clock latency. The special clock topology naturally utilizes setup time slacks to gain hold time margins. Moreover, to alleviate the incurred on-chip variations due to near-threshold computing, we maximize the common clock path shared by flip-flops of each individual stage. Finally, we perform inverter pair swap to maintain duty cycle. Experimental results show that our methodology is promising for industrial bitcoin mining designs: Compared with two variation-aware clock network synthesis approaches widely used in modern ASIC designs, our approach can reduce up to 64% clock buffer/inverter usage, 12% clock power, decrease 99% hold time violating paths, and achieve 85% area saving for timing fixing. The proposed clock design methodology is general and applicable to blockchain and other ASICs with deep pipelines and strong data flow. Chien-Pang Lu, Iris Hui-Ru Jiang, Chih-Wen Yang |
ISPD | 2 |
| 2022 | Deadlock Resolution for Intelligent Intersection Management with Changeable TrajectoriesabstractIntelligent intersection management aims to schedule vehicles so that vehicles can pass through an intersection efficiently and safely. However, inaccurate control, imperfect communication, and malicious information or behavior lead to robustness issues of intelligent intersection management. In this work, we focus on improving robustness against deadlocks by changing the trajectories of vehicles. To guarantee the resolvability of deadlocks, we limit the number of vehicles in an intersection to be smaller than or equal to an intersection-specific value called the maximal deadlock-free load. We develop an algorithm to compute the maximal deadlock-free load. We further reduce the computation time by computing the loads which are pessimistic (smaller) but still deadlock-free. Since the maximal deadlock-free load only depends on the given intersection, it can be integrated with different scheduling algorithms. Experimental results demonstrate that, by changing the trajectories of vehicles and limiting the number of vehicles under maximal deadlock-free loads, our approach can guarantee deadlock-freeness and maintain good traffic efficiency. Li-Heng Lin, Kuan-Chun Wang, Ying-Hua Lee, Kai-En Lin, Chung-Wei Lin, Iris Hui-Ru Jiang |
IV | 6 |
| 2021 | Subresolution Assist Feature Insertion by Variational Adversarial Active Learning and Clustering with Data Point RetrievalabstractAs the feature size keeps shrinking in the modern semiconductor manufacturing process, subresolution assist feature (SRAF) insertion is one promising resolution enhancement technique that can improve the printability and lithographic process window of target patterns. Model-based SRAF generation achieves a high accuracy but with a high computational cost, while rule-based SRAF insertion may require a huge rule table to handle complex patterns. Thus, state-of-the-art works resort to machine learning to reduce runtime but require abundant training samples to generalize the trained models and achieve high performance. Nevertheless, in advanced lithography, we may have a huge solution space of SRAF insertion but few labeled training samples. Therefore, in this work, we address SRAF insertion from a data efficiency perspective. We separate sample selection from SRAF probability learning and train a variational autoencoder and an adversarial network to discriminate between unlabeled and labeled data effectively. Second, we devise a region-based concentric circle area sampling representation to avoid information loss during feature extraction. Third, we determine the final placement of SRAFs by a novel clustering method based on retrieved data points. Experimental results show that, compared with state-of-the-art works, by using 40% training samples, our framework can achieve comparable or even better process variation bands and edge placement errors. Sean Shang-En Tseng, Iris Hui-Ru Jiang, James P. Shiely |
DAC | 2 |
| 2021 | DATC RDF-2021: Design Flow and Beyond ICCAD Special Session PaperabstractThis paper describes the latest release of the DATC Robust Design Flow (RDF), RDF-2021, which has several key additions to expand its horizons. The Chisel/FIRRTL compiler is now part of DATC RDF, enabling support of recent hardware generator designs written in Chisel. Logic locking through RTL obfuscation, an updated ABC synthesis flow, and DFT support are other notable updates to the RDF. A Bookshelf-LEF/DEF converter powered by OpenDB is also added into DATC RDF's inventory as an enabler of robust benchmark conversion. We also describe efforts toward open metrics standards and datasets for machine learning (ML) applications and smart tuning of the design flow, as well as expansion of public analysis calibration data. Our paper closes with future research directions related to DATC's efforts. Jianli Chen, Iris Hui-Ru Jiang, Jinwook Jung, Andrew B. Kahng, Seungwon Kim, Victor N. Kravets, Yih-Lang Li, Ravi Varadarajan, Mingyu Woo |
ICCAD | 2 |
| 2021 | Efficient Mandatory Lane Changing of Connected and Autonomous VehiclesabstractIn a mandatory lane-changing scenario, vehicles need to move to their target lanes before the end of a road segment. Mandatory lane changing usually happens near a highway ramp, and it is one major source of traffic congestion and delay. The advance of connected and autonomous vehicles supports to solve the problem in a centralized and precise way. We formulate a mandatory lane-changing problem and propose a Mixed-Integer Linear Programming (MILP) approach to find an optimal solution. We then propose an MILP-based incremental-window algorithm to solve the problem efficiently. Experimental results show that the MILP-based incremental-window algorithm can take much less computation time and achieve almost the same solution quality, compared with the original MILP approach. The advantage of efficiency yet sufficient effectiveness matches the need of high-level control (decision of passing order) well. Shang-Chien Lin, Chia-Chu Kung, Lee Lin, Chung-Wei Lin, Iris Hui-Ru Jiang |
VTC Fall | 5 |
| 2021 | OpenMPL: An Open-Source Layout DecomposerabstractMultiple patterning lithography has been widely adopted in advanced technology nodes of VLSI manufacturing. As a key step in the design flow, multiple patterning layout decomposition (MPLD) is critical to design closure. Due to the$\mathcal {N} \mathcal {P} $-hardness of the general decomposition problem, various efficient algorithms have been proposed with high-quality solutions. However, with increasingly complicated design flow and peripheral processing steps, developing a high-quality layout decomposer becomes more and more difficult, slowing down further advancement in this field. This article presents$\mathsf {OpenMPL}$(2020), an open-source layout decomposition framework, with well-separated peripheral processing and core solving steps. Besides, previous algorithms or techniques are inspected and several issues are discovered. We then propose corresponding new algorithms to resolve these issues. The experiments demonstrate the effectiveness of our proposed algorithms and the efficiency of$\mathsf {OpenMPL}$. Wei Li 0159, Yuzhe Ma, Qi Sun 0002, Yibo Lin, Iris Hui-Ru Jiang, Bei Yu 0001, David Z. Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2021 | Novel Guiding Template and Mask Assignment for DSA-MP Hybrid Lithography Using Multiple BCP MaterialsabstractDirected self-assembly (DSA) is one of the leading candidates for extending the resolution of optical lithography to sub-7 nm and beyond. By incorporating DSA in multiple patterning lithography (DSA-MP), the flexibility and resolution of contact/via patterning can be further enhanced by using multiple block copolymer (BCP) materials. Prior work faces the dilemma between solution quality and efficiency and is unable to handle 2-D templates. In this article, we capture the essence of template and mask assignment in DSA-MP by a new graph model and a new problem reduction: Our graph model explicitly represents spacing conflict edges and template hyperedges; thus, extra enumeration and manipulation of incompatible via grouping edges can be avoided, and arbitrary 1-D/2-D templates can be natively handled. We further reduce the assignment problem to exact cover, which is encoded by a sparse matrix. Our concise integer linear programming (ILP) formulation and fast backtracking heuristic achieve substantially superior solution quality and efficiency to the state-of-the-art work. Moreover, our problem and graph modelling is flexible and extensible to utilize dummy vias to improve manufacturability. To narrow down the search region of dummy vias, we devise a conflict core finding technique, which is general and applicable to conflict core analysis of exact cover and other multiple patterning layout decomposition problems. Experimental results show the effectiveness and efficiency of our approach. Iris Hui-Ru Jiang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Equivalent Capacitance Guided Dummy Fill Insertion for Timing and ManufacturabilityabstractTo improve manufacturability, dummy fill insertion is widely adopted for reducing the thickness variation after chemical mechanical polishing. However, inserted metal fills induce significant coupling to nearby signal nets, thus possibly incurring timing degradation. Existing timing-aware fill insertion strategies focus on optimizing induced coupling capacitance instead of resultant equivalent capacitance. Therefore, the impact on timing cannot be fully captured. In contrast, in this paper, we analyze equivalent capacitance friendly regions for dummy fills. The analysis can wisely guide dummy fill insertion to prevent unwanted and unnecessary increase in the resultant equivalent capacitance of timing critical nets. Experimental results based on the ICCAD 2018 CAD Contest benchmark suite show that our solution outperforms the contest winning teams and state-of-the-art work. Moreover, our analysis results are highly correlated to actual equivalent capacitance values and indeed provide accurate guidance for timing-aware dummy fill insertion. Sheng-Jung Yu, Chen-Chien Kao, Chia-Han Huang, Iris Hui-Ru Jiang |
ASP-DAC | 4 |
| 2020 | Fast and Accurate Wire Timing Estimation on Tree and Non-Tree Net StructuresabstractTiming optimization is repeatedly performed throughout the entire design flow. The long turn-around time of querying a sign-off timer has become a bottleneck. To break through the bottleneck, a fast and accurate timing estimator is desirable to expedite the pace of timing closure. Unlike gate timing, which is calculated by interpolating lookup tables in cell libraries, wire timing calculation has remained a mystery in timing analysis. The mysterious formula and complex net structures increase the difficulty to correlate with the results generated by a sign-off timer, thus further preventing incremental timing optimization engines from accurate timing estimation without querying a sign-off timer. We attempt to solve the mystery by a novel machine-learning-based wire timing model. Different from prior machine learning models, we first extract topological features to capture the characteristics of RC networks. Then, we propose a loop breaking algorithm to transform non-tree nets into tree structures, and thus non-tree nets can be handled in the same way as tree-structured nets. Experiments are conducted on four industrial designs with tree-like nets (28nm) and two industrial designs with non-tree nets (16nm). Our results show that the prediction model trained by XGBoost is highly accurate: For both tree-like and non-tree nets, the mean error of wire delay is lower than 2 ps. The predicted path arrival times have less than 1% mean error. Experimental results also demonstrate that our model can be trained only once and applied to different designs using the same manufacturing process. Our fast and accurate wire timing prediction can easily be integrated into incremental timing optimization and expedites timing closure. Hsien-Han Cheng, Iris Hui-Ru Jiang, Oscar Ou |
DAC | 2 |
| 2020 | Routing Topology and Time-Division Multiplexing Co-Optimization for Multi-FPGA SystemsabstractTime-division multiplexing (TDM) is widely used to overcome bandwidth limitations and thus enhances routability in multi-FPGA systems due to the shortage of I/O pins in an FPGA. However, multiplexed signals induce significant delays. To evaluate timing degradation, nets with similar criticalities are often grouped to form NetGroups. In this paper, we propose a framework concerning routing topology and time-division multiplexing co-optimization for multi-FPGA systems. The proposed framework first generates high-quality topologies considering Net-Group criticalities. Then, inspired by column generation, TDM ratio assignment is solved optimally by Lagrangian relaxation. Experimental results show that our approach outperforms the top three entries of ICCAD 2019 CAD Contest. Moreover, our TDM ratio assignment algorithm can further improve the results of the top three winners to almost as good as ours. Tung-Wei Lin, Wei-Chen Tai, Iris Hui-Ru Jiang |
DAC | 4 |
| 2020 | Late Breaking Results: Design Dependent Mega Cell Methodology for Area and Power OptimizationabstractTechnology mapping is the key link between technology independent logic synthesis and technology dependent physical design of IC design flow. Conventionally, physical design honors the circuit structure generated by logic synthesis and then performs optimizations to meet the design requirement by using the same cell library as logic synthesis. Thus, the quality of technology mapping is bounded by the variety of library cells. To enhance the flexibility and capability, we propose an analytical mega cell methodology, which clusters the same type or different types of cells together to improve area and power. We analyze the placement of a design and rank mergeable cells for mega cell creation. Through sharing layout space or gate reduction, our approach minimizes the area and power without timing degradation. Our experiments are conducted on five types of SHA256 cores (block chain mining machines) with below 10nm process. Compared with the conventional technology mapping approach (widely adopted by commercial tools), our approach can save average 2.23% area and 14.44% total power consumption. Our results show that the proposed mega cell methodology is promising for energy and area reduction in modern block chain designs and can be easily ported to other ASIC designs. Chien-Pang Lu, Iris Hui-Ru Jiang, Chih-Wen Yang |
DAC | 2 |
| 2020 | DATC RDF-2020: Strengthening the Foundation for Academic Research in IC Physical DesignabstractWe describe the RDF-2020 release of the IEEE CEDA DATC Robust Design Flow (RDF). RDF-2020 extends the previous four years of DATC efforts to (i) preserve and integrate leading research codes, including from past academic contests, and (ii) provide a foundation and backplane for academic research in the RTL-to-GDS IC implementation arena. Implementation and analysis flows have been enhanced by the addition of steps including multi-bit flip-flop clustering, parasitic extraction and antenna checking, as well as a recent contest-winning global router. RDF-2020 also opens a new "Calibrations" direction to support academic research on key analyses such as extraction and timing. An open-source physical design database with Tcl/Python/C++ APIs, a flow integration into a single scriptable application, and support for the newly-opened SKY130 manufacturable PDK, are also new this year. Our paper closes with a discussion of potential future directions for the RDF effort. Jianli Chen, Iris Hui-Ru Jiang, Jinwook Jung, Andrew B. Kahng, Victor N. Kravets, Yih-Lang Li, Shih-Ting Lin, Mingyu Woo |
ICCAD | 2 |
| 2020 | Dynamic IR-Drop ECO Optimization by Cell Movement with Current Waveform Staggering and Machine Learning GuidanceabstractExcessive dynamic IR-drop degrades the circuit performance and may lead to functional failure. Existing IR-drop fixing techniques at the placement stage do not consider the time-variant property and thus cannot handle dynamic IR-drop hotspots well. In current practice, designers perform Engineer Change Order (ECO) to move out these hotspot cells based on their experience. In this paper, we present a novel dynamic IR-drop ECO optimization and prediction framework by wise cell movement. We first spread high demand current cells in a global view to stagger their current waveforms. Then, we further move IR hotspot cells close to power/ground (PG) vias for minimizing the resistance from PG pads to their PG pins. Moreover, we propose an accurate machine learning-based dynamic IR-drop prediction model to guide the final cell movement. The features of our model capture power ground network characteristics, timing information, and cumulative current drawn by cells, thus leading to a general model applicable to ECO. Experimental results show that our proposed model precisely predicts dynamic IR-drop after cell movement, and our optimization scheme can substantially alleviate dynamic IR-drop without timing degradation. Xuan-Xue Huang, Hsien-Chia Chen, Sheng-Wei Wang, Iris Hui-Ru Jiang, Yih-Chih Chou, Cheng-Hong Tsai |
ICCAD | 4 |
| 2020 | Intelligent Design Automation for 2.5/3D Heterogeneous SoC IntegrationabstractAs the design complexity grows dramatically in modern circuit designs, 2.5D/3D chip/package/board integration has become a key to beat process limitation for optimizing system performance and power consumption. Among the explored technologies, the wafer-level integrated fan-out (InFO) package-on-package (PoP) has been adopted by major companies such as TSMC to achieve high-density, high-performance, low-cost packaging solutions. To achieve a high-quality 2.5D/3D heterogeneous integration system, we shall study the chip, package, and board codesign methodology with advanced packages and explore key techniques to handle the emerging challenges in physical design, timing, electrical effects, and testing.1 Iris Hui-Ru Jiang, Yao-Wen Chang, Jiun-Lang Huang, Charlie Chung-Ping Chen |
ICCAD | 1 |
| 2020 | A Dynamic Programming Approach to Optimal Lane Merging of Connected and Autonomous VehiclesabstractLane merging is one of the major sources causing traffic congestion and delay. With the help of vehicle-to-vehicle or vehicle-to-infrastructure communication and autonomous driving technology, there are opportunities to alleviate congestion and delay resulting from lane merging. In this paper, we first summarize modern features and requirements for lane merging, along with the advance of vehicular technology. We then formulate and propose a dynamic programming algorithm to find the optimal solution for a two-lane merging scenario. It schedules the passing order for vehicles while minimizing the time needed for all vehicles to go through the merging point (equivalent to the time that the last vehicle goes through the merging point). We further extend the problem to a consecutive lane-merging scenario. We show the difficulty to apply the original dynamic programming to the consecutive lane-merging scenario and propose an improved version to solve it. Experimental results show that our dynamic programming algorithm can efficiently minimize the time needed for all vehicles to go through the merging point and reduce the average delay of all vehicles, compared with some greedy methods. Shang-Chien Lin, Hsiang Hsu, Chung-Wei Lin, Iris Hui-Ru Jiang, Changliu Liu |
IV | 5 |
| 2020 | iClaire: A Fast and General Layout Pattern Classification Algorithm With Clip Shifting and Centroid RecreationabstractLayout pattern classification, which groups similar layout clips into clusters, underlies a variety of design for manufacturability (DFM) applications, such as hotspot library generation, hierarchical data storage, and yield optimization speedup. The key challenges of layout pattern classification are clip representation and clip clustering. In this paper, we present a fast and general layout pattern classification algorithm considering clip shifting and centroid recreation. Our simple but general clip representation captures both topology and density; we can handle not only rigid area match or edge displacement constraints but also variant edge tolerances and don't care regions. For achieving a small cluster count, our clip clustering is guided by the natural grouping structure of layout clips. The clustering results are further improved by centroid recreation. Our experiments are conducted on 2016 CAD contest at ICCAD benchmark suite. Our results show that our algorithm outperforms the reference solution and all contest winning teams, delivering the smallest cluster count, fastest runtime, and 100% validity. Moreover, our algorithm with clip shifting and centroid recreation further reduces the cluster count effectively and efficiently. In addition to the good solution quality, the interplay between adopted data structures and our algorithm makes it fast and viable to be incorporated into practical DFM flows. Wei-Chun Chang, Iris Hui-Ru Jiang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Novel Guiding Template and Mask Assignment for DSA-MP Hybrid Lithography Using Multiple BCP MaterialsabstractDirected self-assembly (DSA) is one of the leading candidates for extending the resolution of optical lithography to sub-7nm and beyond. By incorporating DSA in multiple patterning lithography (DSA-MP), the flexibility and resolution of contact/via patterning can be further enhanced by using multiple block copolymer (BCP) materials. Prior work faces the dilemma between solution quality and efficiency and is unable to handle 2D templates. In this paper. we capture the essence of template and mask assignment in DSA-MP by a new graph model and a new problem reduction: Our graph model explicitly represents spacing conflict edges and template hyperedges; thus, extra enumeration and manipulation of incompatible via grouping edges can be avoided, and arbitrary 1D/2D templates can be natively handled. We further reduce the assignment problem to exact cover, which is encoded by a sparse matrix. Our concise integer linear programming (ILP) formulation and fast backtracking heuristic achieve substantially superior solution quality and efficiency to the state-of-the-art work. Moreover, our method is flexible and extendible to utilize dummy vias to improve manufacturability. Iris Hui-Ru Jiang |
DAC | 2 |
| 2019 | DATC RDF-2019: Towards a Complete Academic Reference Design FlowabstractWe describe a new RDF-2019 release of the IEEE CEDA DATC Robust Design Flow (RDF). RDF-2019 enhances the DATC RDF to span the entire RTL-to-GDS IC implementation flow, from logic synthesis to detailed routing. The new release represents a significant revision of the previously-reported RDF-2018 flow. Noteworthy vertical extensions include addition of logic synthesis starting from pure behavioral RTL Verilog RTL; floorplanning that includes initial DEF creation, I/O placement and PDN layout generation; and clock tree synthesis between placement legalization and global routing. A number of horizontal extensions to RDF are achieved by incorporating additional tool options at the static timing analysis, global placement, gate sizing, and detailed routing stages of the flow. Further, for the first time, multiple open-source realizations of the entire RDF tool chain are available. Last, RDF-2019 provides significantly enhanced support of and interoperability with industry-standard tools and design formats (LEF/DEF, SPEF, Liberty, SDC, etc.). We illustrate the configuration and use of RDF-2019, with example results on open as well as commercial design enablements. Jianli Chen, Iris Hui-Ru Jiang, Jinwook Jung, Andrew B. Kahng, Victor N. Kravets, Yih-Lang Li, Shih-Ting Lin, Mingyu Woo |
ICCAD | 2 |
| 2019 | Multiple Patterning Layout Compliance with Minimizing Topology Disturbance and Polygon DisplacementabstractMultiple patterning lithography (MPL) divides a layout into several masks and manufactures them by a series of exposure and etching steps. As technology advances, MPL is still indispensable because of its cost effectiveness and hybrid lithography capability. Producing a layout by MPL relies on layout decomposition and layout compliance. The former reports conflicts (i.e., identifies undecomposable polygons), and the latter further modifies the layout to clean conflicts. As long as a layout has unresolved conflicts, it cannot be manufactured by MPL. Hence, layout compliance is crucial for MPL. This task, however, becomes more complicated and challenging because of more masks used and design rule explosion at advanced technology nodes. Semi-automation or manual fixing is thus no longer applicable. Moreover, from a designer's perspective, layout modification is desired to preserve interconnect correctness, not to create new conflicts, and to minimize topology disturbance and polygon displacement. Therefore, in this paper, we propose the first fully automatic approach for multiple patterning layout compliance. For achieving this goal, we extract topology relations of polygons and model the layout correction as a polygon legalization problem. Experimental results demonstrate the superior efficiency and effectiveness of our approach. With topology awareness, our spacing constraint handling is general and can be applied to other layout fixing problems. Hua-Yu Chang, Iris Hui-Ru Jiang |
ISPD | 2 |
| 2019 | Graceful Register Clustering by Effective Mean Shift Algorithm for Power and Timing BalancingabstractAs the wide adoption of FinFET technology in mass production, dynamic power becomes the bottleneck to achieving low power. Therefore, clock power reduction is crucial in modern IC design. Register clustering can effectively save clock power because of significantly reducing the number of clock sinks and register pin capacitance, clock routed wirelength, and the number of clock buffers. In this paper, we propose effective mean shift to naturally form clusters according to register distribution without placement disruption. Effective mean shift fulfills the requirements to be a good register clustering algorithm because it needs no prespecified number of clusters, is insensitive to initializations, is robust to outliers, is tolerant of various register distributions, is efficient and scalable, and balances clock power reduction against timing degradation. Experimental results show that our approach outperforms state-of-the-art work on power and timing balancing, as well as efficiency and scalability. Ya-Chu Chang, Tung-Wei Lin, Iris Hui-Ru Jiang, Gi-Joon Nam |
ISPD | 3 |
| 2019 | Graph-Based Modeling, Scheduling, and Verification for Intersection Management of Intelligent VehiclesabstractIntersection management is one of the most representative applications of intelligent vehicles with connected and autonomous functions. The connectivity provides environmental information that a single vehicle cannot sense, and the autonomy supports precise vehicular control that a human driver cannot achieve. Intersection management solves the fundamental conflict resolution problem for vehicles—two vehicles should not appear at the same location at the same time, and, if they intend to do that, an order should be decided to optimize certain objectives such as the traffic throughput or smoothness. In this paper, we first propose a graph-based model for intersection management. The model is general and applicable to different granularities of intersections and other conflicting scenarios. We then derive formal verification approaches which can guarantee deadlock-freeness. Based on the graph-based model and the verification approaches, we develop a centralized cycle removal algorithm for the graph-based model to schedule vehicles to go through the intersection safely (without collisions) and efficiently without deadlocks. Experimental results demonstrate the expressiveness of the proposed model and the effectiveness and efficiency of the proposed algorithm. Hsiang Hsu, Shang-Chien Lin, Chung-Wei Lin, Iris Hui-Ru Jiang, Changliu Liu |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2018 | FastPass: Fast timing path search for generalized timing exception handlingabstractAs design complexity rapidly grows, a modem design contains more complex constraints and has more clock domains. To these stringent timing requirements, a design is iteratively optimized. Along with intensive optimizations, fast timing analysis guiding designers to fix timing violations is desired. Thus far, previous works have focused on either timing exception handling or path search only. Different from them, in this paper, we tackle these two issues together for the urgent need in modern design. We first generalize timing exceptions to model all common timing exceptions and other path-specific timing quantities. Then, we propose a novel timing analysis flow that performs fast path search for generalized timing exception handling. Furthermore, we develop three delicate techniques to achieve fast path search, including local slack bounds, dynamic slack recovering, and slack priority queue. Experimental results show that our model is general, and our flow is promising with high efficiency and scalability. Pei-Yu Lee, Iris Hui-Ru Jiang, Tung-Chieh Chen |
ASP-DAC | 2 |
| 2018 | COSAT: congestion, obstacle, and slew aware tree construction for multiple power domain designabstractSlew fixing, which ensures correct signal propagation, is essential during timing closure of IC design flow. Conventionally, gate sizing, Vt swapping, or buffer insertion is adopted to locally fix the slew violation on a single gate. Nevertheless, when slew violations are caused by congestion, obstacles, or excessive loadings (e.g., high-fanout nets or long wires), only smart buffering with a global view can fix them. Therefore, in this paper, we propose congestion, obstacle, and slew aware buffered tree construction for excessive loading nets in modern multiple power domain designs. We iteratively cluster sinks into groups by diamond covering and construct Steiner minimal trees. We globally maintain a congestion and obstacle grid map to guide fast grid routing to locate buffers, while avoiding congested regions and obstacles without timing degradation. Our experiments are conducted on seven industrial smartphone designs with TSMC 16/10nm process. Compared with the conventional buffer insertion approach (widely adopted by commercial tools), the minimal chain based approach can reduce 17% buffer count, decrease 14% leakage, and achieve 44% runtime speedup, but incur unwanted timing, design rule, power rule, and routing violations. Our approach can reduce 18% buffer count, decrease 21% leakage, and achieve 92% runtime speedup, while significantly reducing timing, design rule, power rule, and routing short violations. Our results show that our approach is promising for slew fixing on excessive loading nets in modern multiple domain designs. Chien-Pang Lu, Iris Hui-Ru Jiang |
DAC | 2 |
| 2018 | DATC RDF: an academic flow from logic synthesis to detailed routingabstractIn this paper, we present DATC Robust Design Flow (RDF) from logic synthesis to detailed routing. We further include detailed placement and detailed routing tools based on recent EDA research contests. We also demonstrate RDF in a scalable cloud infrastructure. Design methodology and cross-stage optimization research can be conducted via RDF. Jinwook Jung, Iris Hui-Ru Jiang, Jianli Chen, Shih-Ting Lin, Yih-Lang Li, Victor N. Kravets, Gi-Joon Nam |
ICCAD | 2 |
| 2018 | OWARU: Free Space-Aware Timing-Driven Incremental Placement With Critical Path SmoothingabstractThis paper presents an incremental timing-driven placement tool, named OWARU. It optimizes timing critical paths through a free space-aware path smoothing: the gates on such paths are relocated to free spaces around the smoothed paths, while incremental static timing analysis is involved to accurately assess timing changes due to the relocation. OWARU is extended to accommodate gate sizing and layer assignment to demonstrate the effectiveness of unified physical synthesis optimizations and incremental placement. The goal is to show that OWARU is an ideal platform for timing closure at later stages of a physical design flow. OWARU is applied on a set of test circuits from 14-nm high-performance commercial microprocessors, which originally failed in timing closure. On average, the worst slack is improved by 63.6%, which corresponds to 5.0% of the clock period; total negative slack is improved by 69.1%. Jinwook Jung, Gi-Joon Nam, Lakshmi N. Reddy, Iris Hui-Ru Jiang, Youngsoo Shin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | iTimerM: A Compact and Accurate Timing Macro Model for Efficient Hierarchical Timing AnalysisabstractAs designs continue to grow in size and complexity, EDA paradigm shifts from flat to hierarchical timing analysis. In this article, we present compact and accurate timing macro modeling, which is the key to efficient and accurate hierarchical timing analysis. Our goal is to contain only a minimal amount of interface logic in our timing macro model. The main idea is to separate the interface logic into variant and constant timing regions. Then, the variant timing region is reserved for accuracy, while the constant timing region is reduced for compactness. For reducing the constant timing region, we propose anchor pin insertion and deletion by generalizing existing timing graph reduction techniques. Furthermore, we devise a lookup table index selection technique to achieve high model accuracy over the possible operating condition range. Compared with two common models used in industry, extracted timing model and interface logic model, our model has high model accuracy and small model size. Based on the TAU 2016 and 2017 timing macro modeling contest benchmark suites, our results show that our algorithm delivers superior efficiency and accuracy: Hierarchical timing analysis using our model can significantly reduce runtime and memory compared with flat timing analysis on the original design. Moreover, our algorithm outperforms TAU 2016 and 2017 contest winners in model accuracy, model size, model generation performance, and model usage performance. Pei-Yu Lee, Iris Hui-Ru Jiang |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2017 | iClaire: A Fast and General Layout Pattern Classification AlgorithmabstractLayout pattern classification, which groups similar layout clips into clusters, underlies a variety of design for manufacturability (DFM) applications such as hotspot library generation, hierarchical data storage, and yield optimization speedup. The key challenges of layout pattern classification are clip representation and clip clustering, while the mutually conflicting concerns are efficiency and solution quality (in terms of cluster count). In this paper, we present a fast and general layout pattern classification algorithm. Our simple but general clip representation captures both topology and density; we can handle not only rigid area match or edge displacement constraints but also variant edge tolerances and don't care regions. On the other hand, for achieving a small cluster count, our clip clustering is guided by the natural grouping structure of layout clips. Our experiments are conducted on 2016 CAD contest at ICCAD benchmark suite; our results show that our algorithm outperforms the reference solution and all contest winning teams, delivering the smallest cluster count, fastest runtime, and 100% validity. In addition to the good solution quality, the interplay between adopted simple and easily manipulated data structures and our algorithm makes it fast and viable to be incorporated into practical DFM flows. Wei-Chun Chang, Iris Hui-Ru Jiang, Yen-Ting Yu, Wei-Fang Liu |
DAC | 2 |
| 2017 | Power and Area Efficient Hold Time Fixing by Free Metal Segment AllocationabstractHold time fixing ensures correct data synchronization, which is essential and serves as the final step of timing closure for IC design. Conventionally, buffer insertion is adopted to fix hold time violations; buffers, however, induce routing difficulty, increase area utilization, and contribute leakage power. Therefore, in this paper, we propose to fix hold time violations by free metal segment allocation for achieving leakage power efficiency and maintaining utilization for mobile and portable devices. At the final step of timing closure, free metal segments and hold violating nets are both fragmented and scattered over the design. We thus partition a design and perform minimum cost network flow to assign proper free metal segments to hold violating nets. Our experiments are conducted on six industrial smartphone designs with TSMC 16nm process, and our results show that compared with the conventional buffer insertion method, our approach can reduce 37% hold time buffer area, promising for saving leakage power and maintaining area utilization---suited to the final step of timing closure. Wei-Lun Chiu, Iris Hui-Ru Jiang, Chien-Pang Lu, Yu-Tung Chang |
DAC | 2 |
| 2017 | Fast low power rule checking for multiple power domain designabstractPower management via multiple power domains can effectively save power by dynamically turning off idle domains. To control domains of a design, introducing low power intent complicates the physical implementation and verification process. During the physical implementation stage, the optimization or manual ECO could be tedious, and error-prone on power/ground signal connections. Therefore, in this paper, we focus on low power rule checking at the physical implementation stage for multiple power domain design. Existing methods adopt an iterative approach, which identifies one error at a time, thus possibly requiring multiple iterations. Different from them, we propose a fast low power rule checking approach to detect all errors at one time. To do so, we separate all paths into inner-domain and cross-domain paths and extract cross-domain net topology before power rule verification. Based on the global topology, we can verify the correctness of connections and detect all errors at the same time. Experimental results show the effectiveness and efficiency of our approach, achieving 3.62X speedups to detect all errors compared with the iterative approach. Moreover, our approach can identify complicated bugs to facilitate subsequent bug fixing. Chien-Pang Lu, Iris Hui-Ru Jiang |
DATE | 2 |
| 2017 | DATC RDF: Robust design flow database: Invited paperabstractIn this paper, we present DATC Robust Design Flow Database covering the stages from logic synthesis to physical design [1]. Based on this database, design flow and cross-stage optimization research can be conducted via various EDA tools developed from academia. Jinwook Jung, Pei-Yu Lee, Yan-Shiun Wu, Nima Karimpour Darav, Iris Hui-Ru Jiang, Victor N. Kravets, Laleh Behjat, Yih-Lang Li, Gi-Joon Nam |
ICCAD | 5 |
| 2017 | iTimerM: Compact and Accurate Timing Macro Modeling for Efficient Hierarchical Timing AnalysisabstractAs designs continue to grow in size and complexity, EDA paradigm shifts from flat to hierarchical timing analysis. In this paper, we propose compact and accurate timing macro modeling, which is the key to achieve efficient and accurate hierarchical timing analysis. Our macro model tries to contain only a minimal amount of interface logic. For timing graph reduction, we propose anchor pin insertion and deletion by generalizing existing reduction techniques. Furthermore, we devise a lookup table index selection technique to achieve high model accuracy over the possible operating condition range. Compared with two common models used in industry, extracted timing model and interface logic model, our model has high model accuracy and small model size. Based on the TAU 2016 timing contest on macro modeling benchmark suite, our results show that our algorithm delivers superior efficiency and accuracy: Hierarchical timing analysis using our model can significantly reduce runtime and memory compared with flat timing analysis on the original design. Moreover, our algorithm outperforms TAU 2016 contest winner in model accuracy, model size, model usage runtime and memory. Pei-Yu Lee, Iris Hui-Ru Jiang, Ting-You Yang |
ISPD | 2 |
| 2017 | Multiple Patterning Layout Decomposition Considering Complex Coloring Rules and Density BalancingabstractMultiple patterning lithography has been recognized as one of the most promising solutions, in addition to extreme ultraviolet lithography, directed self-assembly, nanoimprint lithography, and electron beam lithography, for advancing the resolution limit of conventional optical lithography. Multiple patterning layout decomposition (MPLD) becomes more challenging as advanced technology introduces complex coloring rules. Existing works model MPLD as a graph coloring problem; nevertheless, when complex coloring rules are considered, layout decomposition can no longer be modeled accurately by graph coloring. Therefore, in this paper, for capturing the essence of layout decomposition with complex coloring rules, we model the MPLD problem as an exact cover problem. We then propose a fast and exact MPLD framework based on augmented dancing links. Our method is flexible and general: it can consider the basic and complex coloring rules simultaneously, can maintain density balancing, and can handle quadruple patterning and beyond. Experimental results show that our approach outperforms state-of-the-art works on reported conflicts and stitches and is promising for handling complex coloring rules and density balancing as well. Iris Hui-Ru Jiang, Hua-Yu Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2016 | Reliability, adaptability and flexibility in timing: Buy a life insurance for your circuitsabstractAt nanometer manufacturing technology nodes, process variations affect circuit performance significantly. In addition, performance deterioration of circuits due to aging effects is also increasing. Consequently, a large timing margin is required to maintain yield. To combat the pessimism and the resulting overdesign, aging analysis with highlevel models, on-chip timing margin monitoring and tuning, and flexible delay models of flip-flops can be deployed. This paper gives an overview of the state of the art of applying these techniques to improve the health of circuits. Ulf Schlichtmann, Masanori Hashimoto, Iris Hui-Ru Jiang, Bing Li 0005 |
ASP-DAC | 3 |
| 2016 | Multiple patterning layout decomposition considering complex coloring rulesabstractMultiple patterning lithography has been recognized as one of the most promising solutions, in addition to extreme ultraviolet lithography, directed self-assembly, nanoimprint lithography, and electron beam lithography, for advancing the resolution limit of conventional optical lithography. Multiple patterning layout decomposition (MPLD) becomes more challenging as advanced technology introduces complex coloring rules. Existing works model MPLD as a graph coloring problem; nevertheless, when complex coloring rules are considered, layout decomposition can no longer be modeled accurately by graph coloring. Therefore, in this paper, for capturing the essence of layout decomposition with complex coloring rules, we model the MPLD problem as an exact cover problem. We then propose a fast and exact MPLD framework based on augmented dancing links. Our method is flexible and general: It can consider the basic and complex coloring rules simultaneously, and it can handle quadruple patterning and beyond. Experimental results show that our approach outperforms state-of-the-art works on reported conflicts and stitches and is promising for handling complex coloring rules as well. Hua-Yu Chang, Iris Hui-Ru Jiang |
DAC | 2 |
| 2016 | Resource-aware functional ECO patch generation
An-Che Cheng, Iris Hui-Ru Jiang, Jing-Yang Jou |
DATE | 2 |
| 2016 | OpenDesign flow database: the infrastructure for VLSI design and design automation researchabstractRecently, there have been a slew of design automation contests and released benchmarks. ISPD place & route contests, DAC placement contests, timing analysis contests at TAU and CAD contests at ICCAD are good examples in the past and more of new contests are planned in the upcoming conferences. These are interesting and important events that stimulate the research of the target problems and advance the cutting edge technologies. Nevertheless, most contests focus only on the point tool problems instead of addressing the design flow or co-optimization among design tools. OpenDesign Flow Database platform is developed to direct attentions to the overall design flow from logic synthesis to physical design optimization [1]. The goals are to provide an academic reference design flow based on past CAD contest results, the database for design benchmarks and point tool libraries, and standard design input/output formats to build a customized design flow by composing point tool libraries. Jinwook Jung, Iris Hui-Ru Jiang, Gi-Joon Nam, Victor N. Kravets, Laleh Behjat, Yih-Lang Li |
ICCAD | 2 |
| 2016 | OWARU: free space-aware timing-driven incremental placementabstractThis paper proposes a powerful new technique called “OWARU”1 that re-places and re-sizes multiple gates simultaneously to improve the most critical paths of a design. In essence, it is an incremental timing-driven placement technique integrated with gate sizing optimization that runs in conjunction with static timing analysis to guarantee a WYSIWYG 2 property. The OWARU technique offers several key advantages over previous techniques such as geometrical path straightening via the Bézier-curve algorithm, free space awareness to guarantee a legal placement solution, and an accurate true timing mode. The Bézier-curve geometric smoothing algorithm is extended with new anchor placement techniques to further improve the path placement. Free space aware placement algorithm is further enhanced with multiple gate optimization. The preliminary results are promising. We applied the OWARU technique at the end of industrial strength physical synthesis optimization on high performance microprocessor designs. The technique was extremely effective in improving the most critical path of the tested designs. On timing critical paths that were not fully closed from the previous physical synthesis optimization, the WS (worst slack) is improved by 5.3% of the total clock period and the TNS (total negative slack) improved by 91.3% on average. Jinwook Jung, Gi-Joon Nam, Lakshmi N. Reddy, Iris Hui-Ru Jiang, Youngsoo Shin |
ICCAD | 4 |
| 2016 | Analytical Clustering Score with Application to Postplacement Register ClusteringabstractCircuit clustering is usually done through discrete optimizations to enable circuit size reduction or design-specific cluster formation. In this article, we are interested in the register-clustering technique for clock-power reduction by leveraging new opportunities introduced by multibit flip-flop (MBFF). Currently, INTEGRA is the only existing postplacement MBFF clustering optimizer with a subquadratic time complexity. However, it severely degrades the wirelength, especially for realistic designs, which may nullify the benefits of MBFF clustering. In contrast, we formulate an analytical clustering score with a nonlinear programming framework, in which the wirelength objective can be seamlessly integrated and the solver has empirical subquadratic time complexity. With the MBFF library, the application of our analytical clustering method achieves comparable clock power to the state-of-the-art techniques, but further reduces the wirelength by about 25%. Even without the MBFF library, we can still achieve 30% clock wirelength reduction. In addition, the proposed method can potentially be integrated into an in-placement MBFF clustering solver and be applied to other problems that require formulating clustering scores in their objective functions. Chang Xu 0005, Guojie Luo, Peixin Li, Yiyu Shi 0001, Iris Hui-Ru Jiang |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2015 | Criticality-dependency-aware timing characterization and analysisabstractFor nanometer design, conventional timing analysis may generate over-optimistic results on criticality-dependent paths. A late arrival time at the data input of a flip-flop lengthens the propagation delay from the clock pin to the data output of this flip-flop, thus degrading the timing margins of paths launching from this flip-flop. To remove the optimism, in this paper, we first propose a simple yet effective triangle model to characterize the criticality-dependency effect. Then, we devise a novel criticality-dependency-aware timing analysis flow, which is seamlessly integrated with the common static timing analysis flow. Experimental results show that our approach can effectively analyze the criticality-dependency effect: Based on the proposed triangle model, we can accurately identify all timing-risky flip-flops and capture the induced timing margin degradation. Yu-Ming Yang, King Ho Tam, Iris Hui-Ru Jiang |
DAC | 3 |
| 2015 | iTimerC 2.0: Fast Incremental Timing and CPPR AnalysisabstractTo achieve timing closure, performance-driven optimizations are repeatedly performed throughout the modern IC design flow. Along with these optimization operations, how to incrementally update timing information efficiently and accurately becomes a crucial task for fast turnaround time. On the other hand, to avoid wasteful over-optimization, clock path pessimism should be removed during timing analysis. In order to provide prompt timing information without over-pessimism during iterative optimizations, in this paper, we aim at fast incremental timing and CPPR analysis. We present two delicate techniques, lazy evaluation and lazy propagation, to avoid redundant updates. Our experiments are conducted on the benchmark suite released by TAU 2015 timing analysis contest. Experimental results show that our timer delivers the best results in terms of accuracy, runtime, and memory over all participating teams. Pei-Yu Lee, Iris Hui-Ru Jiang, Cheng-Ruei Li, Wei-Lun Chiu, Yu-Ming Yang |
ICCAD | 2 |
| 2015 | GasStation: Power and Area Efficient Buffering for Multiple Power Domain DesignabstractPower management, which can be effectively realized by dynamically turning on/off power domains, is a key concern for modern mobile devices. Leakage power in multiple power domain design is usually dominated by nets passing through different power domains because prior work allocates always-on buffers for these feedthrough nets. Unlike prior work, which uses always-on buffers, we create specific power islands with normal buffers in this paper; these power islands gather buffers on feedthrough nets together to share the power and area overhead. The buffer allocation is performed by a wave-propagation based algorithm, named GasStation. Compared with the conventional always-on buffer approach, experimental results show that GasStation can reduce 75% buffer area, 34% buffer count, 80% leakage power, and 8% wirelength on feedthrough nets in eight industrial smart phone designs. Chien-Pang Lu, Iris Hui-Ru Jiang, Chin-Hsiung Hsu |
ICCAD | 2 |
| 2015 | Analytical Clustering Score with Application to Post-Placement Multi-Bit Flip-Flop MergingabstractCircuit clustering is usually done through discrete optimizations, with the purpose of circuit size reduction or design-specific cluster formation. Specifically, we are interested in the multi-bit flip-flop (MBFF) design technique for clock power reduction, where all previous works rely on discrete clustering optimizations. For example, INTEGRA was the only existing post-placement MBFF clustering optimizer with a sub-quadratic time complexity. However, it degrades the wirelength severely, especially for realistic designs, which may cancel out the benefits of MBFF clustering. In this paper we enable the formulation of an analytical clustering score in nonlinear programming, where the wirelength objective can be seamlessly integrated. It has sub-quadratic time complexity, reduces the clock power by about 20% as the state-of-the-art techniques, and further reduces the wirelength by about 25%. In addition, the proposed method is promising to be integrated in an in-placement MBFF clustering solver and be applied in other problems which require formulating the clustering score in the objective function. Chang Xu 0005, Peixin Li, Guojie Luo, Yiyu Shi 0001, Iris Hui-Ru Jiang |
ISPD | 5 |
| 2015 | Machine-Learning-Based Hotspot Detection Using Topological Classification and Critical Feature ExtractionabstractBecause of the widening sub-wavelength lithography gap in advanced fabrication technology, lithography hotspot detection has become an essential task in design for manufacturability. Unlike current state-of-the-art works, which unite pattern matching and machine-learning engines, we fully exploit the strengths of machine learning using novel techniques. By combing topological classification and critical feature extraction, our hotspot detection framework achieves very high accuracy. Furthermore, to speed-up the evaluation, we verify only possible layout clips instead of full-layout scanning. We utilize feedback learning and present redundant clip removal to reduce the false alarm. Experimental results show that the proposed framework is very accurate and demonstrates a rapid training convergence. Moreover, our framework outperforms the 2012 CAD contest at International Conference on ComputerAided Design (ICCAD) winner on accuracy and false alarm. Yen-Ting Yu, Geng-He Lin, Iris Hui-Ru Jiang, Charles C. Chiang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2014 | Functional ECO Using Metal-Configurable Gate-Array Spare CellsabstractMetal-configurable gate-array spare cells, which have versatile functionality, are developed to overcome the inflexibility of standard spare cells used in conventional metal-only engineering change order (ECO). In this paper, we focus on functional ECO optimization using the new type of spare cells to fully exploit its strength. We observe that this functional ECO problem has the nature of dynamic logical and physical costs for selecting spare gate arrays. Unlike existing functional ECO works, which perform technology mapping based on ECO patches, we perform reverse mapping from spare gate arrays to handle these dynamic costs. We devise a spare array relation graph to record geometrical adjacency among spare gate arrays and interleave with the and-inverter network of ECO patches. To avoid redundant traversal and monitor the dynamic costs, we adopt A* search to simultaneously traverse and map between the logical ECO network and the physical spare array relation graph. Hua-Yu Chang, Iris Hui-Ru Jiang, Yao-Wen Chang |
DAC | 2 |
| 2014 | Smart grid load balancing techniques via simultaneous switch/tie-line/wire configurationsabstractFast changing power distribution systems request a dynamic system configuration capability of reacting to volatile consumption demands in an economical way. Load balancing in power distribution systems is an essential technique for smart grid that enables reliable electricity delivery to end customers. This paper is the first work focusing on load balancing using switch reconfiguration, tie-line addition, and wire upgrade simultaneously, while existing works adopt only one of the three techniques to configure the power distribution system. We observe that the new load balancing problem induces a new challenge, dynamic topology rotation, which cannot be handled by existing solutions. To overcome this challenge, we first consider bidirectional power flows and formulate the load balancing problem as a mixed-integer quadratically constrained quadratic program (MIQCQP). To reduce the computational complexity, it is further transformed into a mixed-integer linear program (MILP) without loss of optimality. Experimental results show that, on real power distribution networks, our approach produces optimal solutions that are unlikely to be found in ad-hoc heuristics methods. Iris Hui-Ru Jiang, Gi-Joon Nam, Hua-Yu Chang, Sani R. Nassif, Jerry Hayes |
ICCAD | 1 |
| 2014 | The overview of 2014 CAD contest at ICCADabstractContests and their benchmarks have become an important driving force to push our EDA domain forward in different areas lately, such as ISPD, TAU, DAC contests. The annual CAD Contest in Taiwan has been held for 14 consecutive years and has successfully boosted the EDA research momentum in Taiwan. To encourage better research development on timely and practical EDA problems across all domains, CAD Contest is internationalized since 2012 under the joint sponsorship of the IEEE CEDA and Ministry of Education (MOE) of Taiwan. 2012 CAD Contest attracted 56 teams from 7 regions, while 2013 CAD contest attracted 87 teams from 9 regions. Continuing its great success in 2012 and 2013, 2014 CAD contest attracts 93 teams from 9 regions, including Taiwan, Mainland China, Hong Kong, India, Singapore, US, Canada, Brazil, and Russian Federation. Three contest problems on verification, placement, and mask optimization are announced this year and run by industry experts from Cadence and IBM. Topic chair Chih-Jen Hsu of Cadence Design Systems manages the first contest problem, concentrating on efficiently solving the combinational single-output netlists. The efficiency of solving a single-output function highly depends on the CNF encoding and the SAT solving. For the first problem, contestants are required to explore the best CNF encoding and SAT solver setting to solve the most tests within the shortest runtime. Topic chair Myung-Chul Kim of IBM manages the second problem, focusing on incremental timing-driven placement. Placement, which determines locations of circuit elements, is one of the most crucial steps in the modern IC design flow. For the second problem, contestants are required to perform local refinements on a legal design such that the total slack and worst slack are optimized. Topic chair Rasit O. Topaloglu of IBM manages the third problem, exploring lithography mask optimization. In a circuit layout, densities of polygons within windows of interest may have a large variation across such windows in the rest of the layout. To balance the density of polygons, fills are inserted to make the density of several windows uniform. For the third problem, contestants are required to minimize the density variation with least fills. This session will include three presentations from the contest organizers for these contest problems and an award ceremony. Each contest organizer (topic chair) will present detailed information about the corresponding contest problem, including problem description, benchmarks, and evaluation. Along with the contest, a new set of industrial benchmarks for each contest problem will be released and facilitate scientific evaluations of related research results. We expect that the benchmark suites will further play a key driving force to push the advancement of related research. Moreover, we also expect that the participants will submit their works to the subsequent top conferences to boost related research and also extend the impacts of this contest. Iris Hui-Ru Jiang, Natarajan Viswanathan, Tai-Chen Chen, Jin-Fu Li 0001 |
ICCAD | 1 |
| 2014 | iTimerC: common path pessimism removal using effective reduction methodsabstractStatic timing analysis is a key process to guarantee timing closure for modern IC designs. Nevertheless, fast growing design complexities and increasing on-chip variations complicate this process. To capture more accurate timing performance of a design, common path pessimism removal is prevalent to eliminate artificially induced pessimism in clock paths during timing analysis. To avoid exhaustive exploration on all paths in a design, in this paper, we present a novel timing analysis framework removing common path pessimism based on block-based static timing analysis, timing graph reduction, and dynamic bounding. Experimental results show that the proposed method is highly scalable, especially with short runtimes for large-scale designs. Moreover, our approach outperforms TAU 2014 timing contest winners, generating accurate results and achieving more than 2.13X speedups. Yu-Ming Yang, Iris Hui-Ru Jiang |
ICCAD | 3 |
| 2014 | DRC-based hotspot detection considering edge tolerance and incomplete specificationabstractTo improve the yield in current manufacturing processes, the problematic layout configurations, so-called process-hotspots, should be detected and replaced with yield friendly patterns. A hotspot pattern with edge tolerances and incomplete specifications, where edges may vary in a certain range and any layout configurations may exist in its ambit regions, can sufficiently and generally represent a process-hotspot. This type of hotspots, however, cannot be efficiently or correctly detected by using the state-of-the-art string-matching-based method. In this paper, we present an accurate and efficient DRC-based hotspot detection framework to handle hotspot patterns with edge tolerances and incomplete specifications. Unlike existing DRC-based work, which handles only completely specified patterns, we extract critical design rules to represent all possible topologies of hotspot patterns with edge tolerances and incomplete specifications. We further order these rules to iteratively reduce the search regions of a layout during design rule checking. Then, we apply longest common subsequence and linear scan to locate all hotspots accurately and efficiently. Compared with the state-of-the-art work, experimental results show that our approach can reach promising success rates with significant speedups. Yen-Ting Yu, Iris Hui-Ru Jiang, Charles C. Chiang |
ICCAD | 2 |
| 2014 | PushPull: Short-Path Padding for Timing Error Resilient CircuitsabstractModern IC designs are exposed to a wide range of dynamic variations. Traditionally, a conservative timing guardband is required to guarantee correct operations under the worst-case variation, thus leading to performance degradation. To remove the guardband, resilient circuits are proposed. However, the short-path padding (hold time fixing) problem in resilient circuits is far severer than conventional IC design. Therefore, in this paper, we focus on the short-path padding problem to enable the timing error detection and correction mechanism of resilient circuits. Unlike recent prior work adopts greedy heuristics with a local view, we determine the padding values and locations with a global view. Moreover, we utilize spare cells and a dummy metal to further achieve the derived padding values at physical implementation. Experimental results show that our method is promising to validate timing error-resilient circuits. Yu-Ming Yang, Iris Hui-Ru Jiang, Sung-Ting Ho |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Efficient Coverage-Driven Stimulus Generation Using Simultaneous SAT Solving, with Application to SystemVerilogabstractSystemVerilog provides powerful language constructs for verification, and one of them is the covergroup functional coverage model. This model is designed as a complement to assertion verification, that is, it has the advantage of defining cross-coverage over multiple coverage points. In this article, a coverage-driven verification (CDV) approach is formulated as a simultaneous Boolean satisfiability (SAT) problem that is based on covergroups. The coverage bins defined by the functional model are converted into Conjunction Normal Form (CNF) and then solved together by our proposed simultaneous SAT algorithm PLNSAT to generate stimuli for improving coverage. The basic PLNSAT algorithm is then extended in our second proposed algorithm GPLNSAT, which exploits additional information gleaned from the structure of SystemVerilog covergroups. Compared to generating stimuli separately, the simultaneous SAT approaches can share learned knowledge across each coverage target, thus reducing the overall solving time drastically. Experimental results on a UART circuit and the largest ITC benchmark circuits show that the proposed algorithms can achieve 10.8x speedup on average and outperform state-of-the-art techniques in most of the benchmarks. An-Che Cheng, Chia-Chih Yen, Celina G. Val, Sam Bayless, Alan J. Hu, Iris Hui-Ru Jiang, Jing-Yang Jou |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2013 | Machine-learning-based hotspot detection using topological classification and critical feature extractionabstractBecause of the widening sub-wavelength lithography gap in advanced fabrication technology, lithography hotspot detection has become an essential task in design for manufacturability. Current state-of-the-art works unite pattern matching and machine learning engines. Unlike them, we fully exploit the strengths of machine learning using novel techniques. By combing topological classification and critical feature extraction, our hotspot detection framework achieves very high accuracy. Furthermore, to speed up the evaluation, we verify only possible layout clips instead of full-layout scanning. After detection, we filter hotspots to reduce the false alarm. Experimental results show that the proposed framework is very accurate and demonstrates a rapid training convergence. Moreover, our framework outperforms the 2012 CAD Contest at ICCAD winner on accuracy and false alarm. Yen-Ting Yu, Geng-He Lin, Iris Hui-Ru Jiang, Charles C. Chiang |
DAC | 3 |
| 2013 | The overview of 2013 CAD contest at ICCADabstractContests and their benchmarks have become an important driving force to push our EDA domain forward in different areas lately, such as ISPD, TAU, DAC contests. The annual CAD Contest in Taiwan has been held for 13 consecutive years and has successfully boosted the EDA research momentum in Taiwan. To encourage better research development on timely and practical EDA problems across all domains, CAD Contest is internationalized since 2012 under the joint sponsorship of the IEEE CEDA and Ministry of Education (MOE) of Taiwan. 2012 CAD Contest attracted 56 teams from 7 regions, including USA, Japan, Mainland China, Hong Kong, Korea, Italy, and Taiwan. Continuing its great success in 2012, 2013 CAD contest attracts 87 teams from 9 regions, including USA, Canada, Brazil, India, Russia, Japan, Mainland China, Hong Kong and Taiwan, achieving 55% growth. Three contest problems on technology mapping, placement, and mask optimization are announced this year and run by industry experts from Cadence and IBM. Topic chair Hwei-Tseng Wang of Cadence Design Systems manages the first contest problem, concentrating on technology mapping for macro blocks. The implementation of a digital function is more flexible and powerful as technology advances. Therefore, how to fully utilize and reuse macro blocks in a highly optimized design becomes an important issue. However, it is challenging to identify the boundaries of macro blocks in such complex netlists. For the first problem, contestants are required to map and replace a given design by a set of macro blocks as much as possible. Topic chair Myung-Chul Kim of IBM manages the second problem, focusing on the placement finishing step, detailed placement and legalization. Placement, which determines locations of circuit elements, is one of the most crucial steps in the modern IC design flow. Although there are significant improvements on global placement techniques via recent placement contests, the need for high performance detailed placement continues to grow. For the second problem, contestants are required to perform local refinements on a legal design such that the total wirelength, placement/pin density are optimized. Topic chair Shayak Banerjee of IBM manages the third problem, exploring lithography mask optimization. As technology advances, the printed feature size is smaller than the wavelength of the light shining through the mask. The subwavelength gap causes unwanted shape distortions. To compensate these distortions, mask optimization is performed. For the third problem, contestants are required to find the best mask solution for a given pixelated layout. The best mask solution means least EPE violations and minimum process variations over different corners measured by a provided lithography simulation model. This session will include three presentations from the contest organizers for these contest problems and an award ceremony. Each contest organizer (topic chair) will present detailed information about the corresponding contest problem, including problem description, benchmarks, and evaluation. Along with the contest, a new set of industrial benchmarks for each contest problem will be released and facilitate scientific evaluations of related research results. We expect that the benchmark suites will further play a key driving force to push the advancement of related research. Moreover, we also expect that the participants will submit their works to the subsequent top conferences to boost related research and also extend the impacts of this contest. Iris Hui-Ru Jiang, Zhuo Li 0001, Hwei-Tseng Wang, Natarajan Viswanathan |
ICCAD | 1 |
| 2013 | FF-bond: multi-bit flip-flop bonding at placementabstractClock power contributes a significant portion of chip power in modern IC design. Applying multi-bit flip-flops can effectively reduce clock power. State-of-the-art work performs multi-bit flip-flop clustering at the post-placement stage. However, the solution quality may be limited because the combinational gates are immovable during the clustering process. To overcome the deficiency, in this paper, we propose multi-bit flip-flop bonding at placement. Inspired by ionic bonding in Chemistry, we direct flip-flops to merging friendly locations thus facilitating flip-flop merging. Experimental results show that our algorithm, called FF-Bond, can save 27% clock power on average. Compared with state-of-the-art post-placement multi-bit flip-flop clustering, FF-Bond can further reduce 14% clock power. Chang-Cheng Tsai, Yiyu Shi 0001, Guojie Luo, Iris Hui-Ru Jiang |
ISPD | 4 |
| 2013 | PushPull: short path padding for timing error resilient circuitsabstractModern IC designs are exposed to a wide range of dynamic variations. Traditionally, a conservative timing guardband is required to guarantee correct operations under the worst-case variation, thus leading to performance degradation. To remove the guardband, resilient circuits are proposed. However, the short path padding (hold time fixing) problem in resilient circuits is severer than conventional IC design. Therefore, in this paper, we focus on the short path padding problem to enable the timing error detection and correction mechanism of resilient circuits. Unlike recent prior work adopts greedy heuristics with a local view, we determine the padding values and locations with a global view. Moreover, we propose coarse-grained and fine-grained padding allocation methods to further achieve the derived padding values at physical implementation. Experimental results show that our method is promising to validate timing error resilient circuits. Yu-Ming Yang, Iris Hui-Ru Jiang, Sung-Ting Ho |
ISPD | 2 |
| 2013 | Pulsed-Latch Replacement Using Concurrent Time Borrowing and Clock GatingabstractFlip-flops are the most common form of sequencing elements; however, they have a significantly higher sequencing overhead than latches in terms of delay, power, and area. Hence, pulsed latches are a promising option to reduce power for high-performance circuits. In this paper, to save power and compensate for timing violations, we fully utilize the intrinsic time borrowing property of pulsed latches and consider clock gating during pulsed-latch replacement. Experimental results show that our approach can generate very power efficient results. Chih-Long Chang, Iris Hui-Ru Jiang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | ECO Optimization Using Metal-Configurable Gate-Array Spare CellsabstractDue to the rapidly increasing design complexity in modern IC designs, metal-only engineering change order (ECO) becomes inevitable to achieve design closure with a low respin cost. Traditionally, preplaced redundant standard cells are regarded as spare cells. However, these cells are limited by predefined functionalities and locations, and they always consume leakage power despite their inputs being tied off. To overcome the inflexibility and power overhead, a new type of spare cells, called metal-configurable gate-array spare cells, are introduced. In this paper, we address a new ECO problem, which performs design changes using metal-configurable gate-array spare cells. We first study the properties of this new ECO problem and propose a new cost metric, aliveness, to model the capability of a spare gate array. Based on aliveness and routability, we then develop two ECO optimization frameworks, one for timing ECO and the other for functional ECO. Experimental results show that our approach delivers superior efficiency and effectiveness. Hua-Yu Chang, Iris Hui-Ru Jiang, Yao-Wen Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | Timing ECO optimization using metal-configurable gate-array spare cellsabstractDue to the rapidly increasing design complexity in modern IC designs, metal-only engineering change order (ECO) becomes inevitable to achieve design closure with a low respin cost. Traditionally, preplaced redundant standard cells are regarded as spare cells. However, these cells are limited by predefined functionalities and locations, and they always consume leakage power despite their inputs are tied off. To overcome the inflexibility and power overhead, a new type of spare cells, metal-configurable gate-array spare cells, are considered. Therefore, in this paper, we address a new ECO problem: Timing ECO optimization using metal-configurable gate-array spare cells. We first study the properties for this new ECO problem, propose a new metric, aliveness, to model the capability of a spare gate array, and then develop a timing ECO optimization framework based on aliveness, routability, and timing satisfaction. Experimental results show that our approach delivers superior efficiency and effectiveness. Hua-Yu Chang, Iris Hui-Ru Jiang, Yao-Wen Chang |
DAC | 2 |
| 2012 | Accurate process-hotspot detection using critical design rule extractionabstractIn advanced fabrication technology, the sub-wavelength lithography gap causes unwanted layout distortions. Even if a layout passes design rule checking (DRC), it still might contain process hotspots, which are sensitive to the lithographic process. Hence, process-hotspot detection has become a crucial issue. In this paper, we propose an accurate process-hotspot detection framework. Unlike existing DRC-based works, we extract only critical design rules to express the topological features of hotspot patterns. We adopt a two-stage filtering process to locate all hotspots accurately and efficiently. Compared with state-of-the-art DRC-based works, our results show that our approach can reach 100% success rate with significant speedups. Yen-Ting Yu, Ya-Chung Chan, Subarna Sinha, Iris Hui-Ru Jiang, Charles C. Chiang |
DAC | 4 |
| 2012 | Opening: Introduction to CAD contest at ICCAD 2012: CAD contestabstractContests and their benchmarks have become an important driving force to push our EDA domain forward in different areas lately, such as ISPD, TAU, DAC contests. To encourage better research development on timely and practical EDA problems across all domains, a new international CAD Contest is held this year under the joint sponsorship of the IEEE CEDA and Ministry of Education (MOE) of Taiwan. Three contest problems on functional ECO, placement, and litho hotspot identification are announced this year and run by industry experts from Cadence, IBM and Mentor Graphics. Iris Hui-Ru Jiang, Zhuo Li 0001, Yih-Lang Li |
ICCAD | 1 |
| 2012 | Novel pulsed-latch replacement based on time borrowing and spiral clusteringabstractFlip-flops are the most common form of sequencing elements; however, they have a significantly higher sequencing overhead than latches in terms of delay, power, and area. Hence, pulsed-latches are promising to reduce power for high performance circuits. In this paper, we propose a novel pulsed-latch replacement approach to save power and satisfy timing constraints. We fully utilize the intrinsic time borrowing property of pulsed-latches and develop a spiral clustering method with clock gating consideration. In addition, spiral clustering works well for both rectangular and rectilinear shaped layouts; the latter are popular in modern IC design. Experimental results show that our approach can generate very power efficient results. Chih-Long Chang, Iris Hui-Ru Jiang, Yu-Ming Yang, Evan Y.-W. Tsai, Aki S.-H. Chen |
ISPD | 2 |
| 2012 | Timing ECO Optimization Via Bézier Curve Smoothing and Fixability IdentificationabstractDue to the rapidly increasing design complexity in modern integrated circuit design, more and more timing failures are detected at late stages. Without deferring time-to-market, metal-only engineering change order (ECO) is an economical technique to correct these late-found failures. Typically, a design might need to undergo many ECO runs in design houses; consequently, the usage of spare cells for ECO is of significant importance. In this paper, we aim at timing ECO by using as few spare cells as possible. We observe that a path with good timing is desired to be geometrically smooth. Unlike negative slack and gate delay used in most prior work, we propose a new metric of timing criticality, fixability, by considering the smoothness of timing violating paths. To measure the smoothness of a path, we use the Bézier curve as the golden path. Furthermore, in order to concurrently fix timing violations, we derive a propagation property to divide violating paths into independent segments. Based on Bézier curve smoothing, fixability identification, and the propagation property, we develop an efficient algorithm to fix timing violations. Experimental results show that we can effectively resolve all timing violations with significant speedups over the state-of-the-art works. Hua-Yu Chang, Iris Hui-Ru Jiang, Yao-Wen Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | INTEGRA: Fast Multibit Flip-Flop Clustering for Clock Power SavingabstractClock power is the major contributor to dynamic power for modern integrated circuit design. A conventional single-bit flip-flop cell uses an inverter chain with a high drive strength to drive the clock signal. Clustering several such cells and forming a multibit flip-flop can share the drive strength, dynamic power, and area of the inverter chain, and can even save the clock network power and facilitate the skew control. Hence, in this paper, we focus on postplacement multibit flip-flop clustering to gain these benefits. Utilizing the properties of Manhattan distance and coordinate transformation, we model the problem instance by two interval graphs and use a pair of linear-sized sequences as our representation. Without enumerating all possible combinations, we identify only partial sequences that are necessary to cluster flip-flops, thus leading to an efficient clustering scheme. Moreover, our fast coordinate transformation also makes the execution of our algorithm very efficient. The experiments are conducted on industrial circuits. Our results show that concise representation delivers superior efficiency and effectiveness. Even under timing and placement density constraints, clock power saving via multibit flip-flop clustering can still be substantial at postplacement. Iris Hui-Ru Jiang, Chih-Long Chang, Yu-Ming Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2012 | Reliability-Driven Power/Ground Routing for Analog ICsabstractElectromigration and voltage drop (IR-drop) are two major reliability issues in modern IC design. Electromigration gradually creates permanently open or short circuits due to excessive current densities; IR-drop causes insufficient power supply, thus degrading performance or even inducing functional errors because of nonzero wire resistance. Both types of failure can be triggered by insufficient wire widths. Although expanding the wire width alleviates electromigration and IR-drop, unlimited expansion not only increases the routing cost, but may also be infeasible due to the limited routing resource. In addition, electromigration and IR-drop manifest mainly in the power/ground (P/G) network. Therefore, taking wire widths into consideration is desirable to prevent electromigration and IR-drop at P/G routing. Unlike mature digital IC designs, P/G routing in analog ICs has not yet been well studied. In a conventional design, analog designers manually route P/G networks by implementing greedy strategies. However, the growing scale of analog ICs renders manual routing inefficient, and the greedy strategies may be ineffective when electromigration and IR-drop are considered. This study distances itself from conventional manual design and proposes an automatic analog P/G router that considers electromigration and IR-drops. First, employing transportation formulation, this article constructs an electromigration-aware rectilinear Steiner tree with the minimum routing cost. Second, without changing the solution quality, wires are bundled to release routing space for enhancing routability and relaxing congestion. A wire width extension method is subsequently adopted to reduce wire resistance for IR-drop safety. Compared with high-tech designs, the proposed approach achieves equally optimal solutions for electromigration avoidance, with superior efficiencies. Furthermore, via industrial design, experimental results also show the effectiveness and efficiency of the proposed algorithm for electromigration prevention and IR-drop reduction. Jing-Wei Lin, Tsung-Yi Ho, Iris Hui-Ru Jiang |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2012 | ECOS: Stable Matching Based Metal-Only ECO SynthesisabstractTo ease the time-to-market pressure and save the photomask cost, metal-only ECO realizes the last-minute design changes by revising the photomasks of metal layers only. This task is challenging because the pre-injected spare cells are limited in number and in cell types. Metal-only ECO has to implement these functional and/or timing changes using available spare cells. In this paper, we propose a stable matching based metal-only ECO synthesizer, named ECOS, that can implement the incremental design changes correctly without sacrificing timing and routability. The experiments are conducted on nine industrial testcases. These testcases reflect the real difficulties faced by designers and our results show that ECOS is promising for all of them. Iris Hui-Ru Jiang, Hua-Yu Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | WiT: Optimal Wiring Topology for Electromigration AvoidanceabstractDue to excessive current densities, electromigration (EM) may trigger a permanent open- or short-circuit failure in signal wires or power networks in analog or mixed-signal circuits. As the feature size keeps shrinking, this effect becomes a key reliability concern. Hence, in this paper, we focus on wiring topology generation for avoiding EM at the routing stage. Prior works tended towards heuristics; on the contrary, we first claim this problem belongs to class P instead of class NP-hard. Our breakthrough is, via the proof of the greedy-choice property, we successfully model this problem on a multi-source multi-sink flow network and then solve it by a strongly polynomial time algorithm. Experimental results prove the effectiveness and efficiency of our algorithm. Iris Hui-Ru Jiang, Hua-Yu Chang, Chih-Long Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | Simultaneous functional and timing ECOabstractMetal-only ECO is prevalent at design houses to perform incremental design changes to resolve last found functional and/or timing failures. However, it is hard to perform mixed functional and timing changes manually. Prior endeavors focus on functional or timing ECO alone, but we observe that separating them may fail to fix all timing violations. Consequently, this paper presents the first work to perform simultaneous functional and timing ECO. We use an augmented bipartite graph to model both types of ECO. In addition, through comprehensive constant insertion and bridging, the functional capability of each spare cell is enhanced, thus facilitating spare cell selection. Experimental results show that our simultaneous functional and timing ECO engine can successfully resolve mixed functional and timing ECO that is unsolvable by the sequential scheme. Moreover, our engine outperforms the state-of-the-art works for timing ECO with a 117X speedup, and for functional ECO with 6--15% wirelength reductions. Hua-Yu Chang, Iris Hui-Ru Jiang, Yao-Wen Chang |
DAC | 2 |
| 2011 | Timing ECO optimization via Bézier curve smoothing and fixability identificationabstractDue to the rapidly increasing design complexity in modern IC design, more and more timing failures are detected at late stages. Without deferring time-to-market, metal-only ECO is an economical technique to correct these late-found failures. Typically, a design undergoes many ECO runs in design houses; the usage of spare cells is of significant importance. Hence, in this paper, we aim at timing ECO using the least number of spare cells. We observe that a path with good timing is desired to be geometrically smooth. Different from negative slack and gate delay used in most of prior work, we propose a new metric of timing criticality - fixability - considering the smoothness of critical paths. To measure the smoothness of a path, we use Bézier curve as the golden path. Furthermore, in order to concurrently fix timing violations, we derive the dominance property to divide violated paths into independent segments. Based on Bézier curve smoothing, fixability identification, and the dominance property, we develop an efficient algorithm to fix violations. Compared with the state-of-the-art works, experimental results show that our algorithm not only effectively resolves all timing violations with few spare cells but also achieves 22.8X and 42.6X speedups. Hua-Yu Chang, Iris Hui-Ru Jiang, Yao-Wen Chang |
ICCAD | 2 |
| 2011 | INTEGRA: fast multi-bit flip-flop clustering for clock power saving based on interval graphsabstractClock power is the major contributor to dynamic power for modern IC design. A conventional single-bit flip-flop cell uses an inverter chain with a high drive strength to drive the clock signal. Clustering such cells and forming a multi-bit flip-flop can share the drive strength, dynamic power, and area of the inverter chain, even can save the clock network power and facilitate the skew control. Hence, in this paper, we focus on multi-bit flip-flop clustering at post-placement to gain these benefits. Utilizing the properties of Manhattan distance and coordinate transformation, we model the problem instance by two interval graphs and use a pair of linear-size sequences as our representation. Without enumerating all compatible combinations, we extract only partial sequences that are necessary to cluster flip-flops at a time, thus leading to an efficient clustering scheme. Moreover, our coordinate transformation brings fast operations to execute our algorithm. Experimental results show the superior efficiency and effectiveness of our algorithm. Iris Hui-Ru Jiang, Chih-Long Chang, Yu-Ming Yang, Evan Y.-W. Tsai, Lancer S.-F. Chen |
ISPD | 1 |
| 2010 | Live Demo: ECOS 1.0: A metal-only ECO synthesizerabstractTo ease the time-to-market pressure and save the photomask cost, metal-only ECO realizes the last-minute design changes by revising the photomasks of metal layers only. This task is challenging because the pre-injected spare cells are limited in number and in cell types. We develop a metal-only ECO synthesizer, named ECOS, that automates the incremental design changes without sacrificing timing and routability. Via the live demonstration, visitors can experience the superior performance of ECOS. Iris Hui-Ru Jiang, Hua-Yu Chang |
ISCAS | 1 |
| 2010 | Optimal wiring topology for electromigration avoidance considering multiple layers and obstaclesabstractDue to excessive current densities, electromigration may trigger a permanent open- or short-circuit failure in signal wires or power networks in analog or mixed-signal circuits. As the feature size keeps shrinking, this effect becomes a key reliability concern. Hence, in this paper, we focus on wiring topology generation for avoiding electromigration at the routing stage. Prior works tended towards heuristics; on the contrary, we first claim this problem belongs to class P instead of class NP-hard. Our breakthrough is, via the proof of the greedy-choice property, we successfully model this problem on a multi-source multi-sink flow network and then solve it by a strongly polynomial time algorithm. Experimental results prove the effectiveness and efficiency of our algorithm. Iris Hui-Ru Jiang, Hua-Yu Chang, Chih-Long Chang |
ISPD | 1 |
| 2009 | Matching-based minimum-cost spare cell selection for design changesabstractMetal-only ECO realizes the last-minute design changes by revising the photomasks of metal layers only. This task is challenging because the pre-injected spare cells are limited both in number and in cell types. This paper proposes a matching-based ECO synthesizer, named ECOS, that correctly implements the incremental design changes using the available spare cells as well as tries to reduce the prohibitive photomask cost at the same time. The experiments are conducted on five industrial testcases. ECOS uses less photomask costs to complete design changes for all cases than the direct method that transforms the widely-used hand-editing procedure into an automatic one. Iris Hui-Ru Jiang, Hua-Yu Chang, Liang-Gi Chang, Huang-Bi Hung |
DAC | 1 |
| 2009 | VIFI-CMP: variability-tolerant chip-multiprocessors for throughput and powerabstractThis paper proposes a new architecture of variability-tolerant chip-multiprocessor. To mitigate the impact of process variability on throughput and power, voltage and frequency islands are introduced into chip-multiprocessors. Thus, voltage island frequency island chip-multiprocessors enable per-core scaling on the supply voltage and operating frequency. It can naturally collaborate with dynamic voltage frequency scaling. The process variations are characterized through an analytical model, and are quantified through Monte Carlo analysis. Compared with the design without process variations, when 70 threads are run on a chip of 70 small cores, our results show throughput degradation is 0.06%, while power reduction is 36.27%. Wan-Yu Lee, Iris Hui-Ru Jiang |
ACM Great Lakes Symposium on VLSI | 2 |
| 2009 | POSA: Power-state-aware Buffered Tree ConstructionabstractBuffering without considering power states in multiple supply voltage designs may result in infeasible signals. POSA is the first work to handle this issue. Our buffered tree guarantees feasibility all the times, even when some parts of the design shut down. This feature is one of the key techniques to fulfill power-aware design methodology. Iris Hui-Ru Jiang, Ming-Hua Wu |
ISCAS | 1 |
| 2008 | Power-state-aware buffered tree constructionabstractInterconnect delay and low power are two of the main issues in nano technology. Buffer insertion during routing effectively reduces interconnect delay; power state management and multiple supply voltage significantly lower power consumption. However, buffering without considering power states in multiple supply voltage designs may cause the signal integrity problem. This paper first considers power states into buffered tree construction. Based on a hierarchical approach combined with dynamic programming, we can simultaneously minimize power, satisfy timing constraints and maintain signal integrity. Iris Hui-Ru Jiang, Ming-Hua Wu |
ICCD | 1 |
| 2008 | Configurable rectilinear Steiner tree construction for SoC and nano technologiesabstractThe rectilinear Steiner minimal tree (RSMT) problem is essential in physical design. Moreover, the variant constraints for fabrication issues, including obstacle avoidance, multiple routing layers, layer-specific routing directions, cannot be ignored during RSMT construction for modern SoC and nano technologies. This paper proposes a construction-by-correction approach for obstacle-avoiding preferred direction rectilinear Steiner tree construction. Experimental results show that our algorithm is promising and outperforms the state-of-the-art works. Iris Hui-Ru Jiang, Yen-Ting Yu |
ICCD | 1 |
| 2006 | Reliable crosstalk-driven interconnect optimizationabstractAs technology advances apace, crosstalk becomes a design metric of comparable importance to area and delay. This article focuses mainly on the crosstalk issue, specifically on the impacts of physical design and process variation on crosstalk. While the feature size shrinks below 0.25μ m , the impact of process variation on crosstalk increases rapidly. Hence, a crosstalk insensitive design is desirable in the deep submicron regime. In this article, crosstalk sensitivity is referred to as the influence of process variation on crosstalk in a circuit. We show that the lower bound of crosstalk sensitivity grows quadratically, while that of crosstalk increases linearly. Therefore, designers should also consider crosstalk sensitivity, when optimizing other design objectives such as crosstalk, area, and delay. According to our modeling, these objectives are all in posynomial forms, and thus the multi-objective optimization problem can optimally be solved by Lagrangian relaxation. Experimental results show that our method is effective and efficient. For instance, a circuit of 2856 gates and 5272 wires is optimized using 13-minute runtime and 2.8-MB memory on a Pentium III 1.0 GHz PC with 256-MB memory. In particular, by relaxing Lagrange multipliers to the critical paths, it takes only two iterations for all solutions to converge to the global optimal, which is much more efficient than related previous work. This relaxation scheme provides a key insight into the rapid convergence in Lagrangian relaxation. Iris Hui-Ru Jiang, Song-Ra Pan, Yao-Wen Chang, Jing-Yang Jou |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2004 | Simultaneous floor plan and buffer-block optimizationabstractAs technology advances and the number of interconnections among modules rapidly increases, timing closure, and design convergence are the most important concerns. Hence, it is desirable to consider interconnect optimization as early as possible. Previous work for this issue can be classified into two directions: wire planning and buffer-block planning for interconnect-driven floorplanning. Wire planning for interconnect-driven floorplanning does not consider buffer insertion, and buffer-block planning for interconnect-driven floorplanning cannot overcome the limitation of a bad initial floorplan. In this paper, we first address simultaneous floorplanning and buffer-block planning (i.e., integrating buffer-block planning into floorplanning) for interconnect optimization. We adopt simulated annealing to refine a floorplan so that buffers can be inserted more effectively. In each iteration, we construct a routing tree for each net, allocate buffers for all nets, introduce corresponding buffer blocks into the intermediate floorplan, and invoke Lagrangian relaxation to optimize area and satisfy timing requirements. Further, in order to reduce the problem size, we present supermodule partitioning which partitions modules into supermodules. Experimental results show that our method of integrating buffer-block planning into floorplanning can significantly improve the interconnect delay and reduce the number of buffers needed. Based on a set of MCNC benchmark circuits, our approach achieves an average success rate of 86.1% of nets meeting timing constraints, inserts only 272 buffers on average, and consumes an average extra area of only 0.28% over the given floorplan, compared with the average success rate of 62.6%, 1123 buffers, and extra area of 1.05% resulted from a famous recent work presented at ICCAD'99. Iris Hui-Ru Jiang, Yao-Wen Chang, Jing-Yang Jou, Kai-Yuan Chao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2003 | Simultaneous floorplanning and buffer block planningabstractAs technology advances and the number of interconnections among modules rapidly increases, timing closure and design convergence are the most important concerns. Hence, it is desirable to consider interconnect optimization as early as possible. In this paper, we first address simultaneous floorplanning and buffer block planning (i.e., integrating buffer block planning into floorplanning) for interconnect optimization. Experimental results show that our method can significantly improve the interconnect delay and reduce the number of buffers needed. Iris Hui-Ru Jiang, Yao-Wen Chang, Jing-Yang Jou, Kai-Yuan Chao |
ASP-DAC | 1 |
| 2000 | Optimal reliable crosstalk-driven interconnect optimizationabstractArticle Optimal reliable crosstalk-driven interconnect optimization Share on Authors: Iris Hui-Ru Jiang Department of Electronics Engineering, National Chiao Tung University, Hsinchu 30010, Taiwan Department of Electronics Engineering, National Chiao Tung University, Hsinchu 30010, TaiwanView Profile , Song-Ra Pan Department of Computer and Information Science, National Chiao Tung University, Hsinchu 30010, Taiwan Department of Computer and Information Science, National Chiao Tung University, Hsinchu 30010, TaiwanView Profile , Yao-Wen Chang Department of Computer and Information Science, National Chiao Tung University, Hsinchu 30010, Taiwan Department of Computer and Information Science, National Chiao Tung University, Hsinchu 30010, TaiwanView Profile , Jing-Yang Jou Department of Electronics Engineering, National Chiao Tung University, Hsinchu 30010, Taiwan Department of Electronics Engineering, National Chiao Tung University, Hsinchu 30010, TaiwanView Profile Authors Info & Claims ISPD '00: Proceedings of the 2000 international symposium on Physical designMay 2000 Pages 128–133https://doi.org/10.1145/332357.332388Online:01 May 2000Publication History 6citation191DownloadsMetricsTotal Citations6Total Downloads191Last 12 Months4Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Iris Hui-Ru Jiang, Song-Ra Pan, Yao-Wen Chang, Jing-Yang Jou |
ISPD | 1 |
| 2000 | Crosstalk-driven interconnect optimization by simultaneous gate andwire sizingabstractNoise, as well as area, delay, and power, is one of the most important concerns in the design of deep submicrometer integrated circuits. Currently existing algorithms do not handle simultaneous switching conditions of signals for noise minimization. In this paper, we model not only physical coupling capacitance, but also simultaneous switching behavior for noise optimization. Based on Lagrangian relaxation, we present an algorithm which can optimally solve the simultaneous noise, area, delay, and power optimization problem by sizing circuit components. Our algorithm, with linear memory requirement and linear runtime, is very effective and efficient. For example, for a circuit of 6144 wires and 3512 gates, our algorithm solves the simultaneous optimization problem using only 2.1-MB memory and 19.4-min runtime to achieve the precision of within 1% error on a SUN Spare Ultra-I workstation. Iris Hui-Ru Jiang, Yao-Wen Chang, Jing-Yang Jou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1999 | Hierarchical Floorplan Design on the InternetabstractWith the proliferation of transistor count in VLSI design, more and more design groups try to figure out a way to efficiently combine their designs. The Internet features distributed computing and resource sharing. Consequently, a hierarchical floorplan design can be adequately solved in the Internet environment. In this paper, we address the problem of area minimization floorplan design in the Internet environment. We propose a novel algorithm, RMG algorithm. Taking advantage of the Internet, the RMG algorithm reduces the computing time by shortening the critical path in the floorplan tree. With creating floorplan design in the Internet environment, it can be seen that the Internet has advantages for electronic design automation. Jiann-Horng Lin, Jing-Yang Jou, Iris Hui-Ru Jiang |
ASP-DAC | 3 |
| 1999 | Noise-Constrained Performance Optimization by Simultaneous Gate and Wire Sizing Based on Lagrangian RelaxationabstractNoise, as well as area, delay, and power, is one of the most important concerns in the design of deep sub-micron ICs. Currently existing algorithms do not handle simultaneous switching conditions of signals for noise minimization. In this paper, we model not only physical coupling capacitance, but also simultaneous switching behavior for noise optimization. Based on Lagrangian relaxation, we present an algorithm which can optimally solve the simultaneous noise, area, delay, and power optimization problem by sizing circuit components. Our algorithm, with linear memory requirement overall and linear runtime per iteration, is very effective and efficient. For example, for a circuit of 6144 wires and 3512 gates, our algorithm solves the simultaneous optimization problem using only 1.8 MB memory and 47 minute runtime to achieve the precision of within 1% error on a SUN Sparc Ultra-I workstation. 1 Introduction With decreasing feature sizes, higher clock rates, and increasing interconnect... Iris Hui-Ru Jiang, Jing-Yang Jou, Yao-Wen Chang |
DAC | 1 |
| 1999 | A clustering- and probability-based approach for time-multiplexed FPGA partitioningabstractImproving logic density by time-sharing, time-multiplexed FPGAs (TMFPGAs) has become an important research topic for reconfigurable computing. Due to the precedence and capacity constraints in TMFPGAs, the clustering and partitioning problems for TMFPGAs are different from the traditional ones. We propose a two-phase hierarchical approach to solve the partitioning problem for TMFPGAs. With the precedence and capacity considerations for both phases, the first phase clusters nodes to reduce the problem size, and the second phase applies a probability-based iterative improvement approach to minimize cut cost. Experimental results based on the Xilinx TMFPGA architecture show that our algorithm significantly outperforms previous works. Mango Chia-Tso Chao, Guang-Ming Wu, Iris Hui-Ru Jiang, Yao-Wen Chang |
ICCAD | 3 |
| 1999 | Optimum loading dispersion for high-speed tree-type decision circuitryabstractWith increasing density and capacity due to technology scaling, augmenting data (especially in semiconductor memories) burden selection circuitry with exponentially growing capacitive loads. This tendency violates stringent timing requirements. This work ameliorates the situation for k-stage tree-type decision circuitry. We show that for a k-stage binary decision tree, there always exists an optimum solution such that, after the select-signal arrangement, the worst case loading among select signals equals a lower bound. Our proposed procedure not only provides an optimum solution but also minimizes the loading variance. The worst case loading can be reduced up to nearly k/2 times, thus speeding up and saving power up to W2 times or so for the select signal with the heaviest loading. In contrast, excluding one unit-loading select signal, the empirical variance of the remaining (k-1) signals is always less than 1 instead of diverging. Hence, our approach, for timing-driven layout synthesis, is competent to design high-performance tree-type decision circuitry with more accurate timing and power prediction. In addition, by the presented approach, we can have the alternative of optimizing either for k-stage or for (k-1)-stage, meanwhile possibly minimizing the other. Our algorithm, also, can easily be extended for a general k-stage decision tree with r descendants per node, not restricted to a binary tree; the resultant worst case loading could be quite close to the lower bound and reduced up to nearly k(r-1)/r times. Jie-Hong Roland Jiang, Iris Hui-Ru Jiang |
ICCAD | 2 |