Hung-Ming Chen

dblp:70/3994 · DBLP profile ↗
← Back
105ranked-venue papers
14as first author
22since 2021 · last 2026
0000-0001-8173-3131ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 91 · 10 first-author · 21 since 2021Software engineering, systems software and programming languages · 15 · 4 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Optimizing Multibit Flip-Flop Banking via Agile In-Placement PPA Co-Optimization
abstract
Multibit flip-flop (MBFF) banking has been widely adopted to reduce dynamic power, simplify clock tree structures, and minimize layout area. In our analysis, we reveal that early-stage banking, despite potential initial timing degradation, enables more extensive flip-flop merging, with subsequent placement refinements mitigating timing violations. Experimental results, benchmarked against the first-place winner from the 2024 ICCAD CAD Contest and most recent SOTA, demonstrate that our approach delivers competitive performance while providing enhanced design flexibility and superior power reduction.
Huan-Yuan Chen, Yu-Ruei Lin, Mark Po-Hung Lin, Hung-Ming Chen
DATE4
2026 A PVT-Resilient Subthreshold SRAM-Based In-Memory Computing Accelerator With In-Situ Regulation for Energy-Efficient Spiking Neural Networks
abstract
This paper presents a PVT-resilient, subthreshold SRAM-based computing-in-memory (CIM) macro tailored for energy-efficient spiking neural networks (SNNs). The macro integrates in-situ current sensors and distributed voltage regulators to enable robust large-scale (1024 wordlines, 1304 bitlines and 128 shared neuron cells) subthreshold current-mode CIM, mitigating energy overheads and process-voltage-temperature (PVT) sensitivity. The neuron cells adopt a programmable, memory cell-based firing threshold to enhance neuron robustness against PVT variations. The architecture uses a stride-tick batching schedule to significantly reduce buffer overhead with enhanced input data reuse. Exploiting the high sparsity of SNNs, the proposed system demonstrates significant improvements in energy efficiency and variation tolerance. Fabricated in 28-nm CMOS, the prototype attains 93.64\% accuracy on keyword spotting, delivers up to 1181.42 TOPS/W, and achieves 7.24 TOPS/mm^2, demonstrating a viable and efficient solution for high-performance edge SNN processing.
Shih-Hang Kao, Yang-Chan Hung, I-Wen Wang, Bing-Han Liu, Yu-Chia Chen, Tian-Sheuan Chang, Shyh-Jye Jou, Chien-Nan Jimmy Liu, Hung-Ming Chen, Wei-Zen Chen
IEEE Trans. Circuits Syst. I Regul. Pap.9
2025 On Awareness of Offset-Via and Teardrop in Advanced Packaging Interconnect Synthesis
abstract
In order to take full advantage of chiplet-based system synthesis methodology for HPC and AI applications in high-bandwidth memory, die-to-die designs interconnect need an overhaul breakthrough. The major reason lies in the strengthening technologies: offset-via and teardrop. They need special care in order to enhance reliability and manufacturability. Moreover, conventional signal integrity problems are required to pay attention as well. In this work, by empirical offset-via and layer assignment, we first make sure the optimized routing resources on all the redistribution layers. Then we devise an S-route detailed routing to prevent the detour and to reduce the rip-up and re-route iterations. Results show that we achieved total wire length reduction by 7% on average and total usage of RDLs by nearly 50%, compared with combined SOTA approaches.
Hao-Ju Chang, Yu-Hung Chen, Hao-Wei Huang, Yihua Yeh, Hung-Ming Chen, Chien-Nan Jimmy Liu
ASP-DAC5
2025 Mixed-Size Placement Prototyping Based on Reinforcement Learning with Semi-Concurrent Optimization
abstract
Placement plays a crucial role in modern chip design, aiming to determine the positions of circuit blocks (macros and standard cells). Traditional data structure-centric heuristics often yield suboptimal placement prototypes, ineffectively guiding downstream mixed-size analytical placement to find the desired results for modern large-scale designs. Recent works have showcased the potential of reinforcement learning (RL) to enhance chip placement by training a policy to place macros as a board game. However, placing macros and fixing them in the earlier stages without sufficient information often incurs undesired solutions. This paper proposes a novel RL-based mixed-size placer with iteratively moving the blocks to characterize dense rewards and comprehensive layout information in each step. We further introduce a semi-concurrent moving mechanism to learn the collaborative dynamics among actions on a subset of blocks at each step. We integrate continuous action spaces to develop a deep Q network-based model for learning the semi-concurrent moving policy to derive the proposed moving strategy. Compared with the state-of-the-art methods, experimental results show that our RL-based placer achieves the best placement quality based on commonly used mixed-size placement benchmarks.
Cheng-Yu Chiang, Yi-Hsien Chiang, Chao-Chi Lan, Yang Hsu, Che-Ming Chang, Shao-Chi Huang, Sheng-Hua Wang, Yao-Wen Chang, Hung-Ming Chen
ASP-DAC9
2025 Late Breaking Results: Scalable GPU-Friendly Parallelization for Sweep-Based Maze Routing
abstract
Global routing is a critical stage in the VLSI design flow, aiming to provide a robust guide for detailed routing and serve as early design feedback for placement. Many approaches have leveraged GPU parallelization to achieve significant acceleration. However, with the fast-growing complexity of modern large-scale designs, recent GPU-accelerated maze routing algorithms, driven by the sweep operation, struggle to find solutions efficiently with limited GPU memory resources. In order to address this issue, this paper proposes a scalable, GPU-friendly sweep-based maze routing that requires significantly less memory and fewer kernel function calls while accelerating overall runtime. We introduce a sweep-sharing technique that allows multiple nets to be routed simultaneously within a single sweeping process, substantially reducing memory consumption and kernel launching overhead. We further propose an edge-level rip-up-andreroute technique that selectively reroutes only overflowed segments, preserving feasible parts of the solution to reduce runtime substantially. Experimental results on the latest ISPD’24 Contest benchmarks demonstrate that our GPUfriendly maze routing with sweep sharing can significantly improve the efficiency of the state-of-the-art GPU-accelerated maze router.
Cheng-Yu Chiang, Zong-Ying Cai, Chao-Chi Lan, Yan-Jen Chen, Yang Hsu, Yao-Wen Chang, Hung-Ming Chen
DAC7
2025 Clock and Power Supply-Aware High Accuracy Phase Interpolator Layout Synthesis
abstract
Due to popular requests from the designers of clock and data recovery (CDR) regarding the inefficiency of generating high accuracy phase interpolator (PI), in this work, we have developed a layout generator for such circuit, different from conventional constraint-driven works. In the first stage, we propose a customized template floorplanning plus pin generation demanded by the users. In the second stage, in order to generate high accuracy layout, we implement a gridless router for signal, power supply and clock. Experiments with several configurations indicate that our approach can generate high-quality corresponding layouts that align with user expectations, and even surpass the quality of manual designs on structurally regular high-performance PIs, which are not easy and efficient to be generated by prior primitive/grid-based methods.
Siou-Sian Lin, Shih-Yu Chen, Yu-Ping Huang, Tzu-Chuan Lin, Hung-Ming Chen, Wei-Zen Chen
DATE5
2025 Irregular Operation-Unit-based Compression for Non-Volatile Computation-in-Memory Accelerator
abstract
Irregular pruning significantly enhances the sparsity of deep neural networks (DNNs) and reduces computational demands. However, the irregular sparse weight matrices can diminish the benefits of pruning when implemented in non-volatile computational-in-memory (CIM) accelerators, which relay on dense matrix-vector multiplications. In this work, we propose an irregular operation-unit-based (OU-based) compression method for nonvolatile CIM. We transform the compression problem into a fixed-size clustering task by clustering zero-columns through the rearrangement of matrix rows, followed by their elimination. This approach significantly reduces the number of required non-volatile memory (NVM) macros by up to 2.3x, while compacting the irregular non-zero data for efficient mapping onto non-volatile CIM. We achieve a compression ratio of up to 89% for irregularly pruned weights. Furthermore, we evaluate the mapping of the compacted weights onto nonvolatile CIM accelerator with OU size of 2x128 and 4x128. This work can achieve area efficiency of 0.166 TOPS/mm2and improve area efficiency up to 3.1x compare to state-of-art works.
Liang-Te Huang, Hung-Ming Chen, Po-Tsang Huang
ISCAS3
2025 GIRD: A Green IR-Drop Estimation Method
abstract
An energy-efficient high-performance static IR-drop estimation method based on green learning called Green IR Drop (GIRD) is proposed in this work. GIRD processes the IC design input in three steps. First, the input netlist data are converted to multichannel maps. Their joint spatial–spectral representations are determined with PixelHop. Next, discriminant features are selected using the relevant feature test (RFT). Finally, the selected features are fed to the eXtreme Gradient Boosting trees regressor. Both PixelHop and RFT are green learning tools. GIRD yields a low carbon footprint due to its smaller model sizes and lower computational complexity. Besides, its performance scales well with small training datasets. Experiments on synthetic and real circuits are given to demonstrate the superior performance of GIRD. The model size and the complexity, measured by the floating point operations (FLOPs) of GIRD, are only$10^{-3}$and$10^{-2}$of deep-learning methods, respectively.
Chee-An Yu, Yu-Tung Liu, Yu-Hao Cheng, Shao-Yu Wu, Hung-Ming Chen, C.-C. Jay Kuo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 An Energy-Area-Efficient 3D Interleaved-Ring Accelerator With INT/FP Pipelined PE Array and 3D-SRAM Cube for On-Device CNN Training and DM Inference
abstract
Deep neural networks (DNNs) have demonstrated exceptional performance in image-related artificial intelligence (AI) applications. However, the inference of generative models, such as Diffusion Models (DMs), and the training of convolutional neural networks (CNNs) are computationally intensive tasks, requiring extensive floating-point (FP) operations to maintain accuracy and high-quality results. These tasks are also memory-bound, generating large intermediate data that result in significant external memory access (EMA), thus complicating the deployment of image-related DNNs on edge devices. While 3D-stacked SRAM with through-silicon via (TSV) technology offers promising solutions to alleviate EMA, the complexities of 3D interconnect architectures can introduce substantial overhead in intra-chip communication, potentially degrading overall efficiency. In this paper, we propose a flexible 3D interconnect architecture, termed the 3D interleaved-ring, which utilizes multiple 3D interleaved rings to connect the pipelined integer (INT) and floating-point (FP) processing element (PE) arrays with a 3D-SRAM cube, effectively mitigating the overhead caused by 3D interconnection. We design$3\times 3$micro-routers with dedicated channels and embedded adders for the proposed 3D interconnection architecture. This 3D ring-based architecture reduces the power consumption of on-chip data movement by a factor of 4.5 and decreases the number of required TSVs by a factor of 4.2, compared to conventional 3D mesh-based interconnects. Additionally, we introduce efficient mixed-bit precision dataflows that incorporates dynamic workload distribution to optimize data reuse, reduce bandwidth demands on the 3D-SRAM cube, and improve PE utilization. The proposed work achieves over 90% PE utilization and reduces DRAM accesses by more than$7\times $across various DNN models. Overall, the proposed 3D accelerator improves energy and area efficiency by up to$8.4\times $and$3.4\times $, respectively, compared to state-of-the-art DNN accelerators and processors.
Hung-Ming Chen, Po-Tsang Huang
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 CFIRSTNET: Comprehensive Features for Static IR Drop Estimation with Neural Network
abstract
IR drop estimation is now considered a first-order metric due to the concern about reliability and performance in modern electronic products. Since traditional solution involves lengthy iteration and simulation flow, how to achieve fast yet accurate estimation has become an essential demand. In this work, with the help of modern AI acceleration techniques, we propose a comprehensive solution to combine both the advantages of image-based and netlist-based features in neural network framework and obtain high-quality IR drop prediction very effectively in modern designs. A customized convolutional neural network (CNN) is developed to extract PDN features and make static IR drop estimations. Trained and evaluated with the open-source dataset, experiment results show that we have obtained the best quality in the benchmark on the problem of IR drop estimation in ICCAD CAD Contest 2023, proving the effectiveness of this important design topic.
Yu-Tung Liu, Yu-Hao Cheng, Shao-Yu Wu, Hung-Ming Chen
ICCAD4
2024 A 28nm Energy-Area-Efficient Row-based pipelined Training Accelerator with Mixed FXP4/FP16 for On-Device Transfer Learning
abstract
Training deep convolutional neural networks (DNNs) requires significantly more computational capacity, complex dataflow, memory accesses, and data movement among processing elements (PEs), as well as higher bit precision for back propagation (BP), which demands more power and area overhead than DNN inference. For mobile/edge devices, energy and area efficiency are critical concerns. This research proposes a row-based pipelined DNN training accelerator that employs three techniques to improve energy and area efficiency for resource-constrained edge/mobile devices. The first technique involves freezing weight updates in convolution and batch normalization layers. The second technique involves decomposing the simulated quantization for convolutional layers and reorganizing the operations of batch normalization layers. The mathematical demonstration shows that FP convolution operations can be completed using fixed point (FXP) calculations. FXP MACs with dequantizer can replace the original FP MACs for convolutional layers. Additionally, a row-based FXP/FP pipelined training accelerator is designed for layers pipeline, convolution, and batch normalization layers to increase the FXP and FP resource utilization. The third method uses multi-bank buffer management to prevent data conflicts and reduce the need for on-chip buffers by up to 3.5 times. The proposed accelerator was implemented using the TSMC 28nm CMOS process and achieved an energy efficiency of 2.19 TFLOPS/W and an area efficiency of 85.32 GFLOPS/mm2. It outperforms state-of-the-art works with 6.8 times the area efficiency and 3.7 times the energy efficiency.
Han-Hsiang Pei, Jheng-Rong Yu, Hung-Ming Chen, Po-Tsang Huang
ISCAS4
2024 Enabling System Design in 3D Integration: Technologies and Methodologies
abstract
3D integration solutions have been called for in the semiconductor market for a long time to possibly substitute the place of technology scaling. It consists of 3D IC packaging, 3D IC integration, and 3D silicon integration. 3D IC packaging has been in the market, but 3D IC and silicon integrations have obtained more attention and care due to modern system requirements on high performance computing and edge AI applications. In the need of further integration in electronics system development at lower cost, chip and package design are therefore evolving along the years [11].
Hung-Ming Chen
ISPD1
2023 On Automating Finger-Cap Array Synthesis with Optimal Parasitic Matching for Custom SAR ADC
abstract
Due to its excellent power efficiency, the successive-approximation-register (SAR) analog-to-digital converter (ADC) is an attractive design choice for low-power ADC implements. In analog layout design, the parasitics induced by interconnecting wires and elements affect the accuracy and performance of the device. Due to the requirement of low-power and high-speed, series of very small lateral metal-metal capacitor units are usually adopted as the architecture of capacitor array. Besides power consumption and area reduction, the parasitic capacitance would significantly affect the matching properties and settling time of capacitors. This work presents a framework to synthesize good-quality binary-weighted capacitors for custom SAR ADC. Also, this work proposes a parasitic-aware ILP-based weight-dynamic network routing algorithm to generate a layout considering parasitic capacitance and capacitance ratio mismatch simultaneously. The experimental result shows that the effective number of bits (ENOB) of the layout generated by our approach is comparable to or better than that of manual design and other automated works, closing the gap between pre-sim and post-sim results.
Cheng-Yu Chiang, Chia-Lin Hu, Mark Po-Hung Lin, Yu-Szu Chung, Shyh-Jye Jou, Jieh-Tsorng Wu, Shiuh-Hua Wood Chiang, Chien-Nan Jimmy Liu, Hung-Ming Chen
ASP-DAC9
2023 DPRoute: Deep Learning Framework for Package Routing
abstract
For routing closures in package designs, net order is critical due to complex design rules and severe wire congestion. However, existing solutions are deliberatively designed using heuristics and are difficult to adapt to different design requirements unless updating the algorithm. This work presents a novel deep learning-based routing framework that can keep improving by accumulating data to accommodate increasingly complex design requirements. Based on the initial routing results, we apply deep learning to concurrent detailed routing to deal with the problem of net ordering decisions. We use multi-agent deep reinforcement learning to learn routing schedules between nets. We regard each net as an agent, which needs to consider the actions of other agents while making pathing decisions to avoid routing conflict. Experimental results on industrial package design show that the proposed framework can improve the number of design rule violations by 99.5% and the wirelength by 2.9% for initial routing.
Yeu-Haw Yeh, Simon Yi-Hung Chen, Hung-Ming Chen, Deng-Yao Tu, Guanqi Fang, Yun-Chih Kuo, Po-Yang Chen
ASP-DAC3
2023 Reshaping System Design in 3D Integration: Perspectives and Challenges
abstract
In this paper, we depict modern system design methodologies via 3D integration along with the advance of packaging, considering system prototyping, interconnecting, and physical implementation. The corresponding challenges are presented as well.
Hung-Ming Chen, Chu-Wen Ho, Shih-Hsien Wu, Po-Tsang Huang, Hao-Ju Chang, Chien-Nan Jimmy Liu
ISPD1
2023 Pole-Aware Analog Layout Synthesis Considering Monotonic Current Flows and Wire Crossings
abstract
This article presents a new paradigm for analog placement, which further incorporates poles in addition to the considerations of symmetry island and monotonic current flow while minimizing wire crossings. The nodes along the signal paths in an analog circuit contribute to the poles, and the parasitics on these dominant poles can significantly limit the circuit performance. Although the monotonic placement methods introduced in the previous works can generate simpler routing topologies, the unawareness of poles, especially the dominant and the first non-dominant poles and wire crossings among critical nets, may increase wire load and performance degradation. This article proposes a pole-aware analog layout synthesis methodology to minimize the total wire load and wire crossings while satisfying different symmetry and topological routing constraints. Using a strong-connectivity approach to the group Steiner problem, the presented method for automatic selection of port locations can help reduce total wirelength, increase routing flexibility, and minimize total wire crossings. Experimental results show that the proposed approach results in better solution quality in circuit performance compared with other recent works.
Abhishek Patyal, Hung-Ming Chen, Mark Po-Hung Lin, Guanqi Fang, Simon Yi-Hung Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 On Reducing LDE Variations in Modern Analog Placement
abstract
Layout-dependent (LDEs) introduce an inevitable performance degradation in analog and mixed-signal circuit design with advanced process technologies below 90 nm. The main LDE sources, including the well proximity effect (WPE), length of diffusion (LOD), and the oxide-to-oxide spacing effect (OSE), cause substantial fluctuations in carrier mobility and threshold voltage of transistors. In traditional design flows, impact of these in post-layout simulation, leading to expensive re-design iterations by inspecting the physical locations of devices with respect to one another. In this article, we introduce the concept of an ideal mobility multiplier based on physics models, in order to minimize the LDE effects with a fast simulated annealing algorithm through various LDE alleviating operations. Based on the introduced mobility multiplier and the hierarchical B*-tree (HB*-tree) topological representation, our LDE-aware analog placement methodology can simultaneously optimize not only the area and wire length, but also the LDEs, while maintaining linear-packing time complexity of HB*-trees. Compared to the most recent works on 65 nm-based analog circuits, experimental results show that the proposed method can effectively and efficiently reduce LDE variations, while improving the circuit performance.
A. K. Thasreefa, Abhishek Patyal, Hao-Yu Chi, Mark Po-Hung Lin, Hung-Ming Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2022 Practical Substrate Design Considering Symmetrical and Shielding Routes
abstract
In modern package design, the flip-chip package has become mainstream because of the benefit of its high I/O pins. However, the package design is still done manually in the industry. The lack of automation tools makes the package design cycle longer due to complex routing constraints, and the frequent modification requests. In this work, we propose yet another routing framework for substrate routing. Compared with previous works, our routing algorithm generates a feasible routing solution in a few seconds for industrial design and considers important symmetry and shielding constraints that have not been handled before. Benefiting from the efficiency of our routing algorithm, the designer can get the result immediately and accommodate some modifications to reduce the cost. The experimental result shows that the routing result generated from our router is in good quality, very close to the manual design.
Hao-Yu Chi, Simon Yi-Hung Chen, Hung-Ming Chen, Chien-Nan Jimmy Liu, Yun-Chih Kuo, Ya-Hsin Chang, Kuan-Hsien Ho
DATE3
2022 DASC: A DRAM Data Mapping Methodology for Sparse Convolutional Neural Networks
abstract
The data transferring of sheer model size of CNN (Convolution Neural Network) has become one of the main performance challenges in modern intelligent systems. Although pruning can trim down substantial amount of non-effective neurons, the excessive DRAM accesses of the non-zero data in a sparse network still dominate the overall system performance. Proper data mapping can enable efficient DRAM accesses for a CNN. However, previous DRAM mapping methods focus on dense CNN and become less effective when handling the compressed format and irregular accesses of sparse CNN. The extensive design space search for mapping parameters also results in a time-consuming process. This paper proposes DASC: a DRAM data mapping methodology for sparse CNNs. DASC is designed to handle the data access patterns and block schedule of sparse CNN to attain good spatial locality and efficient DRAM accesses. The bank-group feature in modern DDR is further exploited to enhance processing parallelism. DASC also introduces an analytical model to facilitate fast exploration and quick convergence of parameter search in minutes instead of days from previous work. When compared with the state-of-the-art, DASC decreases the total DRAM latencies and attains an average of 17.1x, 14.3x, and 23.3x better DRAM performance for sparse AlexNet, VGG-16, and ResNet-50 respectively.
Bo-Cheng Lai, Tzu-Chieh Chiang, Po-Shen Kuo, Wan-Ching Wang, Yan-Lin Hung, Hung-Ming Chen, Chien-Nan Jimmy Liu, Shyh-Jye Jou
DATE6
2022 Innovative service model of information services based on the sustainability balanced scorecard: Applied integration of the fuzzy Delphi method, Kano model, and TRIZ
Hung-Ming Chen, Hung-Yi Wu, Pih-Shuw Chen
Expert Syst. Appl.1
2021 Improving the Quality of FPGA RO-PUF by Principal Component Analysis (PCA)
abstract
Ring Oscillator Physical Unclonable Functions (RO-PUFs) exploit the inherent manufacturing process variations, such as systematic and stochastic variations, to generate secret PUF responses that are unique to the device. Stochastic variations are random, while systematic variation exhibits a strong spatial correlation. Therefore, systematic process variation reduces the randomness of the PUF response. This lowers the ability of a PUF response to uniquely identify and authenticate individual devices. Further, the impact of systematic variation is paramount when the two ROs in comparison are placed far apart. Comparing the ROs that are close to each other does improve the randomness, but the responses generated are unreliable and limiting the possible Challenge-Response Pairs (CRPs). In this article, we are proposing a method to reduce the impact of systematic process variation on the RO oscillation frequencies by using Principal Component Analysis (PCA). Principal Components (PCs) model the directions of systematic and stochastic variation present on a device. By projecting the oscillation frequencies in the direction of stochastic variation, the impact of systematic variation can be reduced. Our proposed method neither restricts the placement of ROs to close groups nor limits the possible CRPs. The method is evaluated on a large population of 218 Xilinx Artix-7 FPGAs. To evaluate the efficiency of the proposed method, we purposely paired the ROs that are placed far apart on the FPGA fabric. Results obtained prove the ability of the proposed method in removing the impact of systematic variation on the oscillation frequencies and thereby producing truly random responses.
K. A. Asha, Li En Hsu, Abhishek Patyal, Hung-Ming Chen
ACM J. Emerg. Technol. Comput. Syst.4
2021 A Style-Based Analog Layout Migration Technique With Complete Routing Behavior Preservation
abstract
Layout migration is a fast methodology to generate the required layout for given circuits with different device attributes or different technology. By keeping the original layout topology, the previous design experience can help to keep the circuit performance. However, routing preservation is often not mentioned in previous layout migration techniques, which requires a complete rerouting to break the original style. Panet al.(2015) proposed a topological slicing tree and a constrained delaunay triangulation (CDT) model to keep the routing style during migration. However, this approach may incur some missing nets after migration, which still requires tedious manual works to fix those nets. In this article, we propose a sequence pair (SP)-based placement migration methodology and a novel Cartesian detection line (CDL) model to preserve the routing styles in original layouts. By using the proposed approach, the routability information and routing behaviors can be preserved during layout migration. In order to prevent from missing nets, several refinement techniques are also proposed to fix unreasonable routing nets due to block displacement. In the experiments, the missing nets after migration can be reduced to almost zero with the proposed CDL model, which greatly reduces the extra design efforts.
Hao-Yu Chi, Zi-Jun Lin, Chia-Hao Hung, Chien-Nan Jimmy Liu, Hung-Ming Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2020 Late Breaking Results: Pole-aware Analog Placement Considering Monotonic Current Flow and Crossing-Wire Minimization
abstract
This paper presents a new paradigm for analog placement, which further incorporates poles in addition to the considerations of symmetry-island and monotonic current flow while minimizing wire crossings. The nodes along the signal path in an analog circuit contribute to the poles, and the parasitics on these dominant poles can significantly limit the circuit performance. Although the monotonic placements introduced in the previous works can generate simpler routing topologies, the unawareness of poles, especially both dominant pole and the first non-dominant pole, and wire crossing among critical nets may result in the increase wire-load and performance degradation. Experimental results show that the proposed pole-aware analog placement method considering symmetry-island, monotonic current flow, and crossing-wire minimization results in much better solution quality in terms of circuit performance.
Abhishek Patyal, Hung-Ming Chen, Mark Po-Hung Lin
DAC2
2020 On Pre-Assignment Route Prototyping for Irregular Bumps on BGA Packages
abstract
In modern package design, the bumps often place irregularly due to the macros varied in sizes and positions. This will make pre-assignment routing more difficult, even with massive design efforts. This work presents a 2-stage routing method which can be applied to an arbitrary bump placement on 2-layer BGA packages. Our approach combines escape routing with via assignment: the escape routing is used to handle the irregular bumps and the via assignment is applied for improving the wire congestion and total wirelength of global routing. Experimental results based on industrial cases show that our methodology can solve the routing efficiently, and we have achieved 82% improvement on wire congestion with 5% wirelength increase compared with conventional regular treatments.
Jyun-Ru Jiang, Yun-Chih Kuo, Simon Yi-Hung Chen, Hung-Ming Chen
DATE4
2020 On EDA Solutions for Reconfigurable Memory-Centric AI Edge Applications
abstract
Memory-centric designs deploy computation to storage and enable efficient in-memory computation while avoiding massive amount of data movement. The in-memory-computing schemes have shown distinct advantages and concerns when applying to different types of memory technologies, from conventional SRAM, DRAM to emerging ReRAM. Moreover, the next-generation smart edge systems are expected to support various intelligent applications by employing multi-task machine learning models which would be dynamically activated. To attain an efficient design within short design cycle, it is imperative to have an integrated design framework with automated tools to support hybrid memory systems and perform effective optimization across design stages. This work will introduce a unified framework which integrates EDA solutions to address the design and optimization challenges at different aspects of next-generation memory-centric designs, including fast reconfiguring in-memory/near-memory computing designs to provide optimized solutions (behavioral models and APR cell layouts) for designers to choose the best suitable architectures for their applications.
Hung-Ming Chen, Chia-Lin Hu, Kang-Yu Chang, Alexandra Küster, Yu-Hsien Lin, Po-Shen Kuo, Wei-Tung Chao, Bo-Cheng Lai, Chien-Nan Jimmy Liu, Shyh-Jye Jou
ICCAD1
2020 Achieving Analog Layout Integrity through Learning and Migration Invited Talk
abstract
Analog IC designers and design houses have been accumulating their own design knowledge and constructing their own analog design repositories, including various design specifications, applications, and process technologies. As most of the analog layouts are handcrafted art works, different designers/companies may have different layout guidelines and preferences. When generating a new layout for certain analog design which already exists or is similar to any of those in the repositories, but with different circuit parameters or process technology files, applying layout migration is usually more preferable than starting from scratch. This paper introduces a holistic framework and new layout generation methodology to achieve analog layout integrity through learning and migration. The introduced methodology can effectively and efficiently preserve the preferences of placement and routing topologies from legacy layouts to new ones.
Mark Po-Hung Lin, Hao-Yu Chi, Abhishek Patyal, Zheng-Yao Liu, Jun-Jie Zhao, Chien-Nan Jimmy Liu, Hung-Ming Chen
ICCAD7
2020 Timing Driven Partition for Multi-FPGA Systems with TDM Awareness
abstract
Multi-FPGA system is a popular approach to achieve hardware acceleration with the scalability to accommodate large designs. To overcome the connectivity constraint between each pair of FPGAs, Time-division multiplexing (TDM) is adopted with the expense of additional delay that dominates the performance on multi-FPGA system based emulator. To the best of our knowledge, there is no prior work on partitioning for multi-FPGA system considering hardware configuration and the impact of TDM. This work proposes a partition methodology to improve timing performance for multi-FPGA system. Delay introduced by TDM is estimated and optimized using look-up table for better efficiency. Our experimental result shows 43% improvement in maximum delay while considering both hardware configuration and impact of TDM compared with cut driven partition approach.
Sin-Hong Liou, Sean S.-Y. Liu, Richard Sun, Hung-Ming Chen
ISPD4
2020 Exploring Multiple Analog Placements With Partial-Monotonic Current Paths and Symmetry Constraints Using PCP-SP
abstract
Modern analog placement techniques require consideration of current path and symmetry constraints. The symmetry pairs can be efficiently packed using the symmetry island configurations, but not all these configurations result in minimum gate interconnection, which can impact the overall circuit routing and performance. This article proposes the first work that reformulates this problem considering all of them together in the form of parallel current path (PCP) constraints. PCP constraints, in addition to monotonic current paths, also consider partial-monotonic current paths to generate a more compact placement. We use a novel two-step approach to detect symmetry-feasible sequence-pairs (SFSPs) without doing placement construction by using representative sequence-pair (RSP). Then a placement algorithm satisfying these constraints is formulated to reduce a vast search space via efficient sequence pair manipulation. The experimental results show that this formulation and algorithm can generate multiple placement solutions that satisfy all the constraints in a more tightly packed configuration, resulting in smaller wirelength, reduced parasitics, and thus better post-layout performance.
Abhishek Patyal, Po-Cheng Pan, K. A. Asha, Hung-Ming Chen, Wei-Zen Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 Wire Load Oriented Analog Routing with Matching Constraints
abstract
As design complexity is increased exponentially, electronic design automation (EDA) tools are essential to reduce design efforts. However, the analog layout design has still been done manually for decades because it is a sensitive and error-prone task. Tool-generated layouts are still not well-accepted by analog designers due to the performance loss under non-ideal effects. Most previous works focus on adding more layout constraints on the analog placement. Routing the nets is thus considered as a trivial step that can be done by typical digital routing methodology, which is to use vias to connect every horizontal and vertical lines. Those extra vias will significantly increase the wire loads and degrade the circuit performance. Therefore, in this article, a wire load oriented analog routing methodology is proposed to reduce the number of layer changing of each routing net. Wire load is considered in the optimization goal as well as the wire length to keep the circuit performance after layout, while the analog layout constraints like symmetry and length matching are still satisfied during routing. As shown in the experimental results, this approach significantly reduces the wire load and performance loss after layout with little overhead on wire length.
Hao-Yu Chi, Chien-Nan Jimmy Liu, Hung-Ming Chen
ACM Trans. Design Autom. Electr. Syst.3
2019 An Efficient Learning-based Approach for Performance Exploration on Analog and RF Circuit Synthesis
abstract
An efficient synthesis technique for modern analog circuits is important yet challenging due to the repeatedly re-synthesis process. To precisely explore the analog circuit performance limitation on the required technology is time-consuming. This work presents a learning-based framework for searching the limitation of analog circuits. With hierarchical architecture, the dimension of solution space can be reduced. Bayesian linear regression and support vector machine model are selected to speed up the algorithm and better performance quality can be retrieved. Experimental results show that our approach on two analog circuits can achieve up to 9x runtime speed-up without surrendering performance qualities.
Po-Cheng Pan, Chien-Chia Huang, Hung-Ming Chen
DAC3
2019 Achieving Routing Integrity in Analog Layout Migration via Cartesian Detection Lines
abstract
In order to improve design productivity, proper layout automation tools are desired for analog circuits. Layout migration is one possible approach to generate a new layout for given circuits with different device sizes or different technology, and still keep the original layout topology. However, routing behaviors are often not mentioned in previous works, which requires a complete rerouting that may not follow the original style. Pan [16] first proposed a Constrained Delaunay Triangulation (CDT) based model to keep the routing behavior during layout migration. However, because the device sizes and related distance may be different in the new layout, some reference lines in CDT models may be removed, resulting in some missing nets after migration. In this paper, a novel Cartesian Detection Line (CDL) based model is proposed to preserve the routing behavior in original layouts. Because alternative lines in the modified placement can be easily found to prevent from missing nets, the proposed CDL model greatly improves the routing completeness during layout migration. Several routing refinement techniques are also proposed to solve the routing issues due to block displacement. In our experiments, the routing completeness can be improved to almost 100% with the proposed CDL model, which greatly reduces the design efforts.
Hao-Yu Chi, Zi-Jun Lin, Chia-Hao Hung, Chien-Nan Jimmy Liu, Hung-Ming Chen
ICCAD5
2019 More Effective Power Network Prototyping by Analytical and Centroid Learning
abstract
Recently a prior work has been proposed to improve the power distribution network (PDN) design with some practical methodologies. However, we found that such approach will cause redundant resources, resulting in the waste of the metal application. In this paper, we present a more effective design flow to automatically generate a PDN verified by the commercial tool without IR-Drop violation. We propose an analytical model and consider the different types of macros to determine the total metal width of PDN. Moreover, the optimization is based on a centroid learning method from unsupervised learning to consolidate PDN. Our work has experimented on real designs in 65 nm process, 0.18 um generic process, and 40 nm process. The results show that our framework can satisfy the given IR-Drop constraints and simultaneously save lots of metal resource (means no overdesign).
Yu-Hsiang Chuang, Chang-Tzu Lin, Hung-Ming Chen, Chi-Han Lee, Ting-Sheng Chen
ISCAS3
2018 Performance-preserved analog routing methodology via wire load reduction
abstract
Analog layout automation is a popular research direction in recent years to raise the design productivity. However, the research on this topic is still not well accepted by analog designers because notable performance loss often exists in tool-generated layout. Most previous works focus on layout placement problem and route the nets implicitly by typical digital routing methodology. This routing approach can solve the net crossing issue easily, but requires a lot of extra vias to connect the horizontal and vertical lines, which significantly increases the wire loads and reduces the circuit performance. In the proposed analog routing flow, we try to route each net with minimum layer changing and consider the wire length simultaneously. In other words, wire load is used as the optimization goal instead of using wire length only to keep the circuit performance after laying out the design. As demonstrated on several cases, this approach significantly reduces the wire load and keeps the similar circuit performance as in manual works.
Hao-Yu Chi, Hwa-Yi Tseng, Chien-Nan Jimmy Liu, Hung-Ming Chen
ASP-DAC4
2018 Multi-level droplet routing in active-matrix based digital-microfluidic biochips
abstract
Active-Matrix (AM) technology is currently being used to implement a superior class of EWOD-based biochips, which consist of a dense 2D-array of microelectrodes. These chips offer many advantages over conventional biochips such as the capability of handling variable-size droplets, more flexibility in droplet movement, precise control over droplet navigation, and as a sequel, ease of implementing complex bioprotocols on-chip. However, the new technology poses a number of challenges concerning droplet routing. In order to enhance routability, we propose, in this paper, a multi-level hierarchical approach that takes appropriate decisions on droplet splitting and reshaping. Compared to the most recent routing methods used for EWOD, the proposed multi-level router reduces maximum latest-arrivaltime by an average 18% and achieves 7% less average latest-arrival-time.
Guan-Ruei Lu, Bhargab B. Bhattacharya, Tsung-Yi Ho, Hung-Ming Chen
ASP-DAC4
2018 Analog placement with current flow and symmetry constraints using PCP-SP
abstract
Modern analog placement techniques require consideration of current path and symmetry constraints. The symmetry pairs can be efficiently packed using the symmetry island configurations, but not all these configurations result in minimum gate interconnection, which can impact the overall circuit routing and performance. This paper proposes the first work that reformulates this problem considering all of them together in the form of Parallel Current Path (PCP) constraints. Then a placement algorithm satisfying these constraints is formulated to reduce a vast search space via efficient sequence pair manipulation. Experimental results show that this formulation and algorithm can satisfy all the constraints in a more tightly packed configuration, resulting in lesser routing length, reduced parasitics and thus better post-layout performance.
Abhishek Patyal, Po-Cheng Pan, K. A. Asha, Hung-Ming Chen, Hao-Yu Chi, Chien-Nan Jimmy Liu
DAC4
2018 Extending ML-OARSMT to net open locator with efficient and effective boolean operations
abstract
Multi-layer obstacle-avoiding rectilinear Steiner minimal tree (ML-OARSMT) problem has been extensively studied in recent years. In this work, we consider a variant of ML-OARSMT problem and extend the applicability to the net open location finder. Since ECO or router limitations may cause the open nets, we come up with a framework to detect and reconnect existing nets to resolve the net opens. Different from prior connection graph based approach, we propose a technique by applying efficient Boolean operations to repair net opens. Our method has good quality and scalability and is highly parallelizable. Compared with the results of ICCAD-2017 contest, we show that our proposed algorithm can achieve the smallest cost with 4.81 speedup in average than the top-3 winners.
Bing-Hui Jiang, Hung-Ming Chen
ICCAD2
2018 Reliability Hardening Mechanisms in Cyber-Physical Digital-Microfluidic Biochips
abstract
In the area of biomedical engineering, digital-microfluidic biochips (DMFBs) have received considerable attention because of their capability of providing an efficient and reliable platform for conducting point-of-care clinical diagnostics. System reliability, in turn, mandates error-recoverability while implementing biochemical assays on-chip for medical applications. Unfortunately, the technology of DMFBs is not yet fully equipped to handle error-recovery from various microfluidic operations involving droplet motion and reaction. Recently, a number of cyber-physical systems have been proposed to provide real-time checking and error-recovery in assays based on the feedback received from a few on-chip checkpoints. However, to synthesize robust feedback systems for different types of DMFBs, certain practical issues need to be considered such as co-optimization of checkpoint placement, error-recoverability, and layout of droplet-routing pathways. For application-specific DMFBs, we propose here an algorithm that minimizes the number of checkpoints and determines their locations to cover every path in a given droplet-routing solution. Next, for general-purpose DMFBs, where the checkpoints are pre-deployed in specific locations, we present a checkpoint-aware routing algorithm such that every droplet-routing path passes through at least one checkpoint to enable error-recovery and to ensure physical routability of all droplets. Furthermore, we also propose strategies for executing the algorithms in reliable mode to enhance error-recoverability. The proposed methods thus provide reliability-hardening mechanisms for a wide class of cyber-physical DMFBs.
Guan-Ruei Lu, Ansuman Banerjee, Bhargab B. Bhattacharya, Tsung-Yi Ho, Hung-Ming Chen
ACM J. Emerg. Technol. Comput. Syst.5
2018 Flexible Droplet Routing in Active Matrix-Based Digital Microfluidic Biochips
abstract
The active matrix (AM)-based architecture offers many advantages over conventional digital electrowetting-on-dielectric (EWOD) microfluidic biochips, such as the capability of handling variable-size droplets, more flexible droplet movement, and precise control over droplet navigation. However, a major challenge in choosing the routing paths is to decide when the droplets are to be reshaped depending on the congestion of the intended path, or split- and route sub droplets,and merging them at their respective destinations. As the number of microelectrodes in AM-EWOD chips is large, the path selection problem becomes further complicated. In this article, we propose a negotiation-guided flow based on routing of subdroplets that obviates the explicit need for deciding when the droplets are to be manipulated, yet fully utilizing the power of droplet reshaping, splitting, and merging them to facilitate their journey. The proposed algorithm reduces routing cost and provides more freedom in deadlock avoidance in the presence of multiple routing tasks by assigning certain congestion penalty for sibling subdroplets and fluidic penalty for heterogeneous droplets. Compared to existing techniques, it reduces latest arrival time by an average of 29% for several benchmark and random test suites. Furthermore, our method is observed to provide 100% routability of nets for all test cases, whereas existing and baseline routers fail to produce feasible solutions in many instances. We also propose a reliable mode droplet routing strategy where the number of unreliable splitting operations can be reduced by paying a small penalty on latest arrival time.
Guan-Ruei Lu, Chun-Hao Kuo, Kuen-Cheng Chiang, Ansuman Banerjee, Bhargab B. Bhattacharya, Tsung-Yi Ho, Hung-Ming Chen
ACM Trans. Design Autom. Electr. Syst.7
2017 Heterogeneous chip power delivery modeling and co-synthesis for practical 3DIC realization
abstract
Three dimensional IC (3DIC) is becoming practical in today's consumer electronics designs. However, one major problem remains in design synthesis and flow: how to model heterogeneous die(s) with major logic die for power synthesis and signoff. This work provides a realistic model and principle for heterogeneous dies power network for 3DICs. It is based on given abstract or early stage information like bump location and power consumption from the provider. Our work also uses this model to synthesize power network with bottom logic die in the design flow. The result is DRC clean power network without IR and EM violation for all power domains. First, we analyze the location and power consumption of power bump for heterogeneous die(s). Second, according to previous analysis, we decide the stripe location and power sink location of heterogeneous dies model by a clustering method. After the initial model is synthesized, we convert it to a node graph with corresponding resistance of via and metal layer, also nodal voltages. Third, the model is optimized by using Sequential Linear Programming (SLP) to adjust stripe width. It will improve the model iteratively until the target IR-Drop is met. Furthermore, our work will create a pseudo DEF of the proposed model to be incorporated with the commercial tool for verification. We experiment on a real case from design house containing a 3D DRAM stack to demonstrate the effectiveness of this cross-layer realization. Results show that we can save 34% metal layer usage in one of the power domains in our case by using proposed methodology.
Wei-Hsun Liao, Chang-Tzu Lin, Sheng-Hsin Fang, Chien-Chia Huang, Hung-Ming Chen, Ding-Ming Kwai, Yung-Fa Chou
ASP-DAC5
2017 On reliability hardening in cyber-physical digital-microfluidic biochips
abstract
In the area of biomedical engineering, digital-microfluidic biochips (DMFBs) have received considerable attention, because of their capability of providing an efficient and reliable platform for conducting point-of-care clinical diagnostics. System reliability, in turn, mandates error-recoverability while implementing biochemical assays on-chip for medical applications. Unfortunately, the technology of DMFBs is not yet fully equipped to handle error-recovery from various microfluidic operations involving droplet motion and reaction. Recently, a number of cyber-physical systems have been proposed to provide real-time checking and error-recovery in assays based on the feedback received from a few on-chip checkpoints. However, in order to synthesize robust feedback systems for different types of DMFBs, certain practical issues need to be considered such as co-optimization of checkpoint placement and layout of droplet-routing pathways. For application-specific DMFBs, we propose here an algorithm that minimizes the number of checkpoints and determines their locations to cover every path in a given droplet-routing solution. Next, for general-purpose DMFBs, where the checkpoints are pre-deployed in specific locations, we present a checkpoint-aware routing algorithm such that every droplet-routing path passes through at least one checkpoint to enable error-recovery and to ensure physical routability of all droplets. Our experiments on assay benchmarks show encouraging results in terms of latest-arrival-time and routability of droplets. The proposed methods thus provide convenient reliability-hardening mechanisms for a wide class of cyber-physical DMFBs.
Guan-Ruei Lu, Guan-Ming Huang, Ansuman Banerjee, Bhargab B. Bhattacharya, Tsung-Yi Ho, Hung-Ming Chen
ASP-DAC6
2016 A New Methodology for Noise Sensor Placement Based on Association Rule Mining
abstract
Due to near-threshold computing nowadays, voltage emergency is threatening our design margins very seriously. Noise sensors are inserted in order to prevent various integrity issues from happening during runtime. In this work, we use a new technique based on association rule mining to plan and place noise sensors. This new methodology can consider the miss rate (the probability of any node occurring voltage emergency without any detection by placed sensors) and simultaneously minimize the number of sensors utilized.The results show that our approach is very effective in converging the miss rate to zero by the least number of sensors. Compared with the state-of-the-art, we can reduce the number of sensors by half in benchmarks while the miss rate is comparable or even smaller than the prior work.
Yu-Hsiang Hung, Sheng-Hsin Fang, Hung-Ming Chen, Shen-Min Chen, Chang-Tzu Lin, Chia-Hsin Lee
ACM Great Lakes Symposium on VLSI3
2016 Using a branch-and-bound and a genetic algorithm for a single-machine total late work scheduling problem
Chin-Chia Wu, Yunqiang Yin, Wen-Hsiang Wu, Hung-Ming Chen, Shuenn-Ren Cheng
Soft Comput.4
2015 An approach to anchoring and placing high performance custom digital designs
abstract
Custom layouts of digital blocks are often used in mixed-signal designs in order to meet the critical performance requirements. Unlike traditional standard-cell based digital placement, custom-cell based digital placement may need additional manual help and intervention to achieve higher performance. The need for manual intervention is primarily due to the inability of modern analytical placers in delivering satisfactory performance on placing designs without pre-placed blocks. While most design flow works in a flat or top-down fashion, custom digital design generally works in a bottom-up fashion that there is no prior knowledge on I/O pin plan since it is changeable by the owners of modules. Without any or very few fixed I/O locations, modern analytical placers tend to produce unsatisfactory results. In this work, we propose a method, mimicking the process of making beds, to guide state-of-the-art analytical placers to deliver better placement results. With the crafted pseudo anchors and nets, total HPWL on Capo10.5 [1], mPL6 [2], NTUplace3 [3] and VDAPlace [4] have improved by 2.92%, 8.69%, 25.26% and 10.72% respectively on a set of industry custom digital designs and improved by 2.19%, 4.34%, 36.09%, and 14.27% respectively on Peko-Suite1 benchmarks.
Shih-Ying Liu, Tung-Chieh Chen, Hung-Ming Chen
ASP-DAC3
2015 Closing the Gap between Global and Detailed Placement: Techniques for Improving Routability
abstract
Improving routability during both global and detailed routing stage has become a critical problem in modern VLSI design. In this work, we propose a placement framework that offers a complete coverage solution in considering both global and detailed routing congestion. A placement migration strategy is proposed, which improves detailed routing congestion while preserving the placement integrity that is optimized for global routability. Using the benchmarks released from ISPD2014 Contest, practical design rules in advanced node design are considered in our placement framework. Evaluation on routability of our placement framework is conducted using commercial router provided by the 2014 ISPD Contest organizers. Experimental results show that the proposed methodologies can effectively improve placement solutions for both global and detailed router.
Chun-Kai Wang, Chuan-Chia Huang, Shih-Ying Liu, Ching-Yu Chin, Sheng-Te Hu, Wei-Chen Wu, Hung-Ming Chen
ISPD7
2015 A Fast Prototyping Framework for Analog Layout Migration With Planar Preservation
abstract
Analog layout generation in the advanced CMOS design is challenging by its increasing layout constraints and performance requirements. This situation becomes more intricate by the growing parasitic variability and manufacturing reliability. To facilitate the feasibility of template-based layout migration, this paper first introduces a layout preservation, which extracts placement and routing behaviors from an existing layout into a crossing graph via constrained Delaunay triangulation. And later this crossing graph can be migrated into multiple layouts with placement and routing reconnection. The proposed approach also provides a refinement for wire to optimize the performance metrics. This approach is applied to a variable-gain amplifier, a folded-cascode operational amplifier, and a low dropout regulator. The experimental results demonstrate more possibility on layout migration, such that averagely more than 75% routing of migrated layout is generated by our approach. Additionally, it exhibits the productivity with qualified performance on different designs.
Po-Cheng Pan, Ching-Yu Chin, Hung-Ming Chen, Tung-Chieh Chen, Chin-Chieh Lee, Jou-Chun Lin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2014 Routability-driven bump assignment for chip-package co-design
abstract
In current chip and package designs, it is a bottleneck to simultaneously optimize both pin assignment and pin routing for different design domains (chip, package, and board). Usually the whole process costs a huge manual effort and multiple iterations thus reducing profit margin. Therefore, we propose a fast heuristic chip-package co-design algorithm in order to automatically obtain a bump assignment which introduces high routability both in RDL routing and substrate routing (100% in our real case). Experimental results show that the proposed method (inspired by board escape routing algorithms) automatically finishes bump assignment, RDL routing and substrate routing in a short time, while the traditional co-design flow requires weeks even months.
Meng-Ling Chen, Tu-Hsiung Tsai, Hung-Ming Chen, Shi-Hao Chen
ASP-DAC3
2014 Cost-effective decap selection for beyond die power integrity
abstract
In designing reliable power distribution networks (PDN) for power integrity (PI), it is essential to stabilize voltage supply to devices on chip. We usually employ decoupling capacitor (decap) to suppress the noise generated by the switching of devices. There have been numerous prior works on how to select/insert decaps in chip, package, or board to maintain PI, however optimal decap selection is usually not applicable due to design budget and manufacturability. Moreover, design cost is seldom touched or mentioned. In this research, we propose an efficient methodology “PDC-PSO” to automatically optimizing the selection of available decaps. This algorithm not only takes advantage of particle swarm optimization (PSO) to stochastically search the design space, but takes the most effective range of decaps into consideration to outperform the basic PSO. We apply this to three real package designs and the results show that, compared to the original decap selection by rules of thumb, our approach could shorten the design period and we have better combination of decaps at the same or lower cost. In addition, our methodology can also consider package-board co-design in optimizing different operation frequencies.
Yi-En Chen, Tu-Hsiung Tsai, Shi-Hao Chen, Hung-Ming Chen
DATE4
2014 Memcomputing: The cape of good hope: [Extended special session description]
abstract
Energy efficiency has emerged as a major barrier to performance scalability for modern processors. On the other hand, significant breakthroughs have been achieved in memory technologies recently [1-4, 6]. As such, the fascinating idea of memcomputing (i.e., use memory for computation purposes) has drawn wide attention from both academia and industry as an effective remedy. Compared with conventional logic computing, memory array provides large set of parallel resources with high bandwidth, which can be configured to perform in-situ computing and information processing, leading to drastic reduction in processor-memory traffic. It will not only make computations more power-and speed-efficient, but also smarter. In addition, it exploits the advances in memory technologies (e.g., [8, 9]) and integration approaches (e.g. 3D integration [11-17]) to achieve better technology scalability. This special session includes three presentations that offer a broad-spectrum retreat on this hot topic.
Yiyu Shi 0001, Hung-Ming Chen
DATE2
2014 Planning and placing power clamps for effective CDM protection
abstract
The issue on reliability of the device becomes more critical as power density of device progressively increases with advancement of technology nodes. Smaller transistor and hence thinner gate oxide implies transistors are more vulnerable against an Electrostatic Discharge (ESD) event. Among the test models in ESD, Charged Device Model (CDM) has greater potential to cause catastrophic damage to the device due to its faster and larger discharging current. To protect against a CDM event, power clamps are placed across the design to offer a low resistance discharge path. However, conventional power clamp placement method to place power clamps generally relies on design experience. In this work, we propose a power clamp placement algorithm that places power clamp at strategic location which can effectively minimize number of power clamps while achieving better protection against a CDM event compared to conventional approach.
Hsin-Chun Lin, Shih-Ying Liu, Hung-Ming Chen
ICCAD3
2014 Improving power delivery network design by practical methodologies
abstract
There are many works on the power network design and prototyping for digital designs, however some usual and practical design concerns are not addressed. In this work, we present a realistic power network design methodology without IR violation certified by state-of-the-art commercial tool. Our work integrates analysis, optimization and synthesis of power network. In particular, we consider thermal effect and power pad's positions during the prototyping of power network. A scenario in placement regarding the violation of design rules is considered and resolved by maximum flow algorithm at the same stage. After the synthesis of initial power network, we generate a sensitivity matrix which is correlated with nodal voltage and resistances of net and via in metal layers. Furthermore, a Sequential Linear Programming(SLP) will be applied to adjust the sensitivity matrix iteratively until the IR drop constraint is satisfied. Our work is experimented on a real design in TSMC 65nm LP process, and the result validates our framework that the IR-Drop can be reduced to 2% of supply voltage.
Chia-Chi Huang, Chang-Tzu Lin, Wei-Syun Liao, Chieh-Jui Lee, Hung-Ming Chen, Chia-Hsin Lee, Ding-Ming Kwai
ICCD5
2014 Automatic image segmentation and classification based on direction texton technique for hemolytic anemia in thin blood smears
Hung-Ming Chen, Ya-Ting Tsao, Shin-Ni Tsai
Mach. Vis. Appl.1
2014 ACER: An Agglomerative Clustering Based Electrode Addressing and Routing Algorithm for Pin-Constrained EWOD Chips
abstract
The problem of pin-constrained electrowetting-ondielectric (EWOD) biochips becomes a serious issue to realize complex bio-chemical operations. Due to limited number of control pins and routing resources, additional Printed Circuit Board (PCB) routing layers may be required which potentially raises the fabrication cost. Previous state-of-the-art work has tried to develop a framework that uses a network-flow-based method for broadcast electrodeaddressing EWOD biochips. Nevertheless, greedily merging of electrical pins in previous works is at high risk of producing unroutable design. Routability should have higher priority than pin reduction. While previous works dedicated their effort on pin reduction, we have addressed our attention on routability of broadcast addressing. Experimental results demonstrate that taking routability into consideration can even have higher pin reduction. Viewed in this light, we present ACER, a routability driven clustering algorithm followed by escape routing using integer linear programming that effectively solves both pin merging and routing in broadcast addressing framework. Our proposed algorithm does not greedily focus on pin-reduction. Instead, routability is taken into consideration through agglomerative clustering. Compared to previous state-of-the-art, our proposed algorithm can further reduce required control pins by an average of 13% and route the design using 68% less wirelength.
Shih-Ying Liu, Chung-Hung Chang, Hung-Ming Chen, Tsung-Yi Ho
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2014 Clock Tree Synthesis Considering Slew Effect on Supply Voltage Variation
abstract
This work tackles a problem of clock power minimization within a skew constraint under supply voltage variation. This problem is defined in the ISPD 2010 benchmark. Unlike mesh and cross link that reduce clock skew uncertainty by multiple driving paths, our focus is on controlling skew uncertainty in the structure of the tree. We observe that slow slew amplifies supply voltage variation, which induces larger path delay variation and skew uncertainty. To obtain the optimality, we formulate a symmetric clock tree synthesis as a mathematical programming problem in which the slew effect is considered by an NLDM-like cell delay variation model. A symmetry-to-asymmetry tree transformation is proposed to further reduce wire loading. Experimental results show that the proposed four methods save up to 20% of clock tree capacitance loading. Beyond controlling slew to suppress supply-voltage-variation-induced skew, we also discuss the strategies of clock tree synthesis under variant variation scenarios and the limitations of the ISPD 2010 benchmark.
Chun-Kai Wang, Yeh-Chi Chang, Hung-Ming Chen, Ching-Yu Chin
ACM Trans. Design Autom. Electr. Syst.3
2014 Fast Thermal Aware Placement With Accurate Thermal Analysis Based on Green Function
abstract
In this paper, we present a fast and accurate thermal aware analytical placer. A thermal model is constructed based on a Green function with discrete cosine transform (DCT) to generate full chip temperature profile. Our thermal model is tightly integrated with an analytical placer implemented based on the SimPL framework. A temperature spreading force based on the Gaussian model is proposed to reduce the maximum on-chip temperature and optimize tradeoff between total half-perimeter wirelength and on-chip maximum temperature. The temperature profile generated using our thermal model is verified by the ANSYS ICEPAK and obtains an average deviation within 3.0% with 240× speedup.
Shih-Ying Liu, Ren-Guo Luo, Suradeth Aroonsantidecha, Ching-Yu Chin, Hung-Ming Chen
IEEE Trans. Very Large Scale Integr. Syst.5
2013 A network-flow based algorithm for power density mitigation at post-placement stage
abstract
In this paper, we propose a power density mitigation algorithm at post-placement stage. Our proposed framework first identifies cluster of bins with high temperature, then propagates power density away from high temperature region by balancing regional power density. The problem of balancing regional power density is modeled as a supply-demand problem and solution is obtained with minimal displacement of cells. An analytical temperature profiling algorithm is tightly integrated within the framework to constantly update the temperature profile in response to incremental perturbation to placement. Our proposed approach can effectively reduce maximum temperature compared to previous works on temperature mitigation.
Shih-Ying Liu, Ren-Guo Luo, Hung-Ming Chen
DATE3
2013 Effective power network prototyping via statistical-based clustering and sequential linear programming
abstract
In this paper, we propose a framework that automatically generates a power network based on given placed design and verifies the power network by the commercial tool without IR and Electro-Migration (EM) violations. Our framework integrates synthesis, optimization and analysis of power network. A deterministic method is proposed to decide number and location of power stripes based on clustering analysis. After an initial power network is synthesized, we propose a sensitivity matrix Gswhich is the correlation between updates in stripe resistance and nodal voltage. An optimization scheme based on Sequential Linear Programming (SLP) is applied to iteratively adjust power network to satisfy a given IR drop constraint. The proposed framework constantly updates voltage distribution in response to incremental change in power network. To accurately capture voltage distribution on a given chip, our power network models every existing power stripes and via resistances on each layer. Experimental result demonstrates that our power network analysis can accurately capture voltage distribution on a given chip and effectively minimize power network area. The proposed methodology is experimented on two real designs in TSMC 90nm and UMC 90nm technology respectively and achieves 9%–32% reduction in power network area, compared with the results from modern commercial PG synthesizer.
Shih-Ying Liu, Chieh-Jui Lee, Chuan-Chia Huang, Hung-Ming Chen, Chang-Tzu Lin, Chia-Hsin Lee
DATE4
2013 PAGE: parallel agile genetic exploration towards utmost performance for analog circuit design
abstract
This paper presents an agile hierarchical synthesis framework for analog circuit. To acknowledge the limitation for a given topology analog circuit, this hierarchical synthesis work proposes a performance exploration technique and a non-uniform-step simulation process. Apart from spec targeted designs, this proposed approach can help to search the solutions better than designers' expectation. A parallel genetic algorithm (PAGE) method is employed for performance exploration. Unlike other evolution-based topology explorations, this is the first method that regards performance constraints as input genome for evolution and resolves the multiple-objective problem with the multiple-population feature. Populations of selected performance are transfered to device variables by re-targeting technique. Based on a normalization of device variable distribution, a probabilistic stochastic simulation significantly reduces the convergence time to find the global optima of circuit performance. Experimental results show that our approach on radio-frequency distributed amplifier (RFDA) and folded cascode operational amplifier (Op-Amp) in different technologies can obtain better runtime and higher quality in analog synthesis.
Po-Cheng Pan, Hung-Ming Chen, Chien-Chih Lin
DATE2
2013 Efficient analog layout prototyping by layout reuse with routing preservation
abstract
To strive for better circuit performance on analog design, layout generation heavily relies on experienced analog designers' effort. Other than general analog constraints such as symmetry and wire-matching are commonly embraced in many proposed works, analog circuit performance is also sensitive to routing behavior. This paper presents a CDT-based layout extraction to preserve routing behavior of the reference layout. Furthermore, a generalized layout prototyping methodology is proposed based on the layout extraction to achieve routing reuse. The proposed layout prototyping is applied to a variable-gain amplifier and a folded-cascode operational amplifier for both migration and prototypes generation. Experimental results show that our approach effectively reduces design cycle time and simultaneously produces reasonable performance.
Ching-Yu Chin, Po-Cheng Pan, Hung-Ming Chen, Tung-Chieh Chen, Jou-Chun Lin
ICCAD3
2013 On the way to practical tools for beyond die codesign and integration
abstract
Package and board designs were usually considered inferior parts in semiconductor supply chain, compared with major stream digital and AMS/RF chip designs. This thought has been gradually reverted due to profit margin reduction in lack of consideration for those ``beyond die"parts. In recent years, Prof. Kajitani has devoted himself in this particular part of researchestowards his retirement and beyond (we are honored to be invited to work together), he has generated numerous usefulthoughts (legacy of object coding) in board routing automation. Although not aware of all his inventions in this field in Japan, we have been sure that Prof. Kajitani's brilliant thoughts have influenced many researchers and students, including us in Taiwan. During his several visits to Taiwan, he was invited to give talks to express his interests in non-maze routing and other more topics. Besides well-known sequence pair representation in floorplanning and placement advances, he has made other influential contributions in beyond die codesign and integration. The ultimate goal of Prof. Kajitani (and our research group) is to try to generate practical tools for non-die layout design automation, and it has been unsurprisingly uneasy task. In this talk, we try to reveal some of the development paths and challenges we have encountered, also showing the records in this cross-nation collaboration towarding this still-continuing mission to possible practicality. We also hope this talk can somewhat help pave the road to the real success of beyond die design automation, regardless of our do-the-best attempt and very little outcome so far.
Hung-Ming Chen
ISPD1
2013 A 3D visualized expert system for maintenance and management of existing building facilities using reliability-based method
Hung-Ming Chen, Chuan-Chien Hou, Yu-Hsiang Wang
Expert Syst. Appl.1
2013 Package routability- and IR-drop-aware finger/pad planning for single chip and stacking IC designs
Chao-Hung Lu, Hung-Ming Chen, Chien-Nan Jimmy Liu, Wen-Yu Shih
Integr.2
2013 Escaped Boundary Pins Routing for High-Speed Boards
abstract
Routing for high-speed boards is still achieved manually today. There have recently been some related works to solve this problem; however, a more practical problem has not been addressed. Usually, the packages or components are designed with or without the requirement from board designers, and the boundary pins are usually fixed or advised to follow when the board design starts. In this paper, we describe this fixed ordering boundary pin routing problem, and propose a practical approach to solve it. Not only do we provide a way to address, we also further plan the wires in a better way to preserve the precious routing resources in the limited number of layers on the board, and to effectively deal with obstacles. Our approach has different features compared with the conventional shortest-path-based routing paradigm. In addition, we consider length-matching requirements and wire shape resemblance for high-speed signal routes on board. Our results show that we can utilize routing resources very carefully, and can account for the resemblance of nets in the presence of the obstacles. Our approach is workable for board buses as well.
Ching-Yu Chin, Chung-Yi Kuan, Tsung-Ying Tsai, Hung-Ming Chen, Yoji Kajitani
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2013 A study of row-based area-array I/O design planning in concurrent chip-package design flow
abstract
IC-centric design flow has been a common paradigm when designing and optimizing a system. Package and board/system designs are usually followed by almost-ready chip designs, which causes long turn-around time communicating with package and system houses. In this article, the realizations of area-array I/O design methodologies are studied. Different from IC-centric flow, we propose a chip-package concurrent design flow to speed up the design time. Along with the flow, we design the I/O-bump (and P/G-bump) tile that combines I/O (and P/G) and bump into a hard macro with the considerations of I/O power connection and electrostatic discharge (ESD) protection. We then employ an I/O-row based scheme to place I/O-bump tiles with existed metal layers. By such a scheme, it reduces efforts in I/O placement legalization and the redistribution layer (RDL) routing. With the emphasis on package design awareness, the proposed methods map package balls onto chip I/Os, thus providing an opportunity to design chip and package in parallel. Due to this early study of I/O and bump planning, faster convergence can be expected with concurrent design flow. The results are encouraging and the merits of this flow are reassuring.
Ren-Jie Lee, Hung-Ming Chen
ACM Trans. Design Autom. Electr. Syst.2
2013 Agglomerative-based flip-flop merging and relocation for signal wirelength and clock tree optimization
abstract
In this article, we propose a flip-flop merging algorithm based on agglomerative clustering. Compared to previous state-of-the-art on flip-flop merging, our proposed algorithm outperforms that of Chang et al. [2010] and Wang et al. [2011] in all aspects, including number of flip-flop reductions, increase in signal wirelength, displacement of flip-flops, and execution time. Our proposed algorithm also has minimal disruption to original placement. In comparison with Jiang et al. [2011], Wang et al. [2011], and Chang et al. [2010], our proposed algorithm has the least displacement when relocating merged flip-flops. While previous works on flip-flop merging focus on the number of flip-flop reduction, we further evaluate the power consumption of clock tree after flip-flop merging. To further minimize clock tree wirelength, we propose a framework that determines a preferable location for relocated merged flip-flops for clock tree synthesis (CTS). Experimental results show that our CTS-driven flip-flop merging can reduce clock tree wirelength by an average of 7.82% with minimum clock network power consumption compared to all of the previous works.
Shih-Ying Liu, Wan-Ting Lo, Chieh-Jui Lee, Hung-Ming Chen
ACM Trans. Design Autom. Electr. Syst.4
2013 Board- and Chip-Aware Package Wire Planning
abstract
The slow turnaround between design, package, and system houses has been one of the primary concerns in the semiconductor business. There is a serious lag in the development time of the systems due to time-consuming interface design between the chip, package, and board. In order to enable chip-package-board codesign to speed up the design process, we propose an approach to address this issue by efficiently planning wires for board and chip design awareness, which includes the package pin-out designation and the corresponding wire planning in package and board. We model the problem as an interval intersection problem. Because of the special need in pin-out rules, an algorithm to resolve the problem is developed. We then use some optimization techniques to further improve objectives such as global wire congestion and length deviation. Our results show that a very efficient estimation can be made considering those important objectives, and package congestion can be successfully mitigated.
Ren-Jie Lee, Hsin-Wu Hsu, Hung-Ming Chen
IEEE Trans. Very Large Scale Integr. Syst.3
2012 A fast thermal aware placement with accurate thermal analysis based on Green function
abstract
In this paper, we propose a fast and accurate thermal aware analytical placer. Thermal model is constructed based on Green function with enhanced DCT to generate full chip temperature profile. Unlike other previous thermal aware placers, our thermal model is tightly integrated with a flat force directed placement. A thermal spreading force based on 2D Gaussian model is proposed to reduce maximum on-chip temperature with dynamic hot region size control, optimizing between total half-perimeter wirelength (HPWL) and on-chip temperature distribution. Our thermal model is evaluated by the most recent commercial tool and has an average deviation of 6.5% with 242× speed up. Our placer can reach the same quality compared to Capo and APlace2 with 2–3× speed up. Experiments are tested using ISPD 2005 benchmark with up to 2 million gate design. The results are further evaluated using GSRC Bookshelf Evaluator for total HPWL, and using ICEPAK for temperature distribution. To the best of our knowledge, this is the first thermal-aware placer using analytical thermal model and experimented on large scale design. It takes 6.5 hours to complete entire 2005 ISPD benchmark using our thermal aware placer.
Suradeth Aroonsantidecha, Shih-Ying Liu, Ching-Yu Chin, Hung-Ming Chen
ASP-DAC4
2012 On effective flip-chip routing via pseudo single redistribution layer
abstract
Due to the advantage of flip-chip design in power distribution but controversial peripheral IO placement in lower design cost, redistribution layer (RDL) is usually used for such interconnection. Sometimes RDL is so congested that the capacity for routing is insufficient. Routing therefore cannot be completed within a single layer even for manual routing. Although [2] proposed a routing algorithm that uses two layers of RDLs, but in practice the required routing area is a little more than one layer. We overcome this problem by adopting the concept of pseudo single-layer. With the heuristics for routing on mapped channels and observations on staggered pins to relieve vertical constraints, the area of 2-layer routing can be minimized and the routability is 100%. Comparisons of routing results between manual design, the commercial tool, and the proposed method are presented. We have shown the effectiveness on a real industrial case: it originally required fully manual design, the proposed method can finish RDL routing automatically and effectively.
Hsin-Wu Hsu, Meng-Ling Chen, Hung-Ming Chen, Hung-Chun Li, Shi-Hao Chen
DATE3
2012 Agglomerative-based flip-flop merging with signal wirelength optimization
abstract
In this paper, an optimization methodology using agglomerative-based clustering for number of flip-flop reduction and signal wirelength minimization is proposed. Comparing to previous works on flip-flop reduction, our method can obtain an optimal tradeoff curve between flip-flop number reduction and increase in signal wirelength. Our proposed methodology outperforms [1] and [12] in both reducing number of flip-flops and minimizing increase in signal wirelength. In comparison with [9], our methodology obtains a tradeoff of 15.8% reduction in flip-flop's signal wirelength with 16.9% additional flip-flops. Due to the nature of agglomerative clustering, when relocating flip-flops, our proposed method minimizes total displacement by an average of 5.9%, 8.0%, 181.4% in comparison with [12], [1] and [9] respectively.
Shih-Ying Liu, Chieh-Jui Lee, Hung-Ming Chen
DATE3
2012 Configurable analog routing methodology via technology and design constraint unification
abstract
In this paper, we present a novel configurable analog routing methodology for more efficient analog layout automation. By the help of OpenAccess constraint group format, the technology process rules and analog layout design intention/constraints are unified through schematic level to layout level. In contrast to self-defined constraint format in prior arts, proposed approach manipulates the analog routing characteristic based on the unified constraints. In different circuit hierarchies defined by circuit designers or extracted by existing placement, the hierarchical structure is formed as specific analog layout constraint groups. This work efficiently facilitates analog routing strategy which honors the specific analog constraints. By practicing on an analog functional block of tsmc 40nm SoC design which guarantees to be legalized and satisfies required analog constraints by DRC/LVS and post-layout simulation respectively, the results in wire matching for signal integrity show that the different routing priority generated by our approach can have significant performance impact.
Po-Cheng Pan, Hung-Ming Chen, Yi-Kan Cheng, Jill Liu, Wei-Yi Hu
ICCAD2
2012 On construction low power and robust clock tree via slew budgeting
abstract
Clock skew resulted by process variation becomes more and more serious as technology shrinks. In 2010, ISPD held a high performance clock network synthesis contest; it considered supply-voltage variation and wire manufacturing variation. Previous works show that the main issue of variation induced skew is on supply-voltage variation. To trade off power and supply-voltage variation induced skew more effectively, we adapt a tree topology which use a timing model independent symmetrical tree at top level to drive the bottom level non-symmetry trees. Our method gives top tree more power budget to reduce supply-voltage variation induced skew and greedily saves power consuming in bottom level. Experimental results are evaluated from the benchmarks of ISPD contest 2010. Compared with state-of-the-art cross link work, the proposed technique reduces 10% of power consumption on average and also improves the run time.
Yeh-Chi Chang, Chun-Kai Wang, Hung-Ming Chen
ISPD3
2012 Genetic programming for predicting aseismic abilities of school buildings
Hung-Ming Chen, Wei-Ko Kao, Hsing-Chih Tsai
Eng. Appl. Artif. Intell.1
2011 Row-based area-array I/O design planning in concurrent chip-package design flow
abstract
IC-centric design flow has been a common paradigm when designing and optimizing a system. Package and board/system designs are usually followed by almost-ready chip designs, which causes long turn-around time communicating with package and system houses. In this paper, the realizations of area-array I/O design methodologies are studied. Different from IC-centric flow, we propose a chip-package concurrent design flow to speed up the design time. Along with the flow, we design the I/O-bump (and P/G-bump) tile which combines I/O (and P/G) and bump into a hard macro with the considerations of I/O power connection and electrostatic discharge (ESD) protection. We then employ an I/O-row based scheme to place I/O-bump tiles with existed metal layers. By such a scheme, it reduces efforts in I/O placement legalization and the redistribution layer (RDL) routing. With the emphasis on package design awareness, the proposed methods map package balls onto chip I/Os, thus providing an opportunity to design chip and package in parallel. Due to this early study of I/O and bump planning, faster convergence can be expected with concurrent design flow. The results are encouraging and the merits of this flow are reassuring.
Ren-Jie Lee, Hung-Ming Chen
ASP-DAC2
2011 On routing fixed escaped boundary pins for high speed boards
abstract
Routing for high speed boards is still achieved manually nowadays. There have been some related works in escape routing to solve this problem recently, however a more practical problem is not addressed. Usually the packages/components are designed with or without the requirement from board designers, and the boundary pins are usually fixed or advised to follow when the board design starts. Previous works in escape routing are not likely to be used due to this nature, in this work, we describe this fixed ordering boundary pin escaping problem, and propose a practical approach to solve it. Not only can we have a way to address, we also further plan the wires in a better way to preserve the precious routing resources in the limited number of layers on the board, and to effectively deal with obstacles. Our approach has different feature compared with conventional shortest-path-based routing paradigm. In addition, we consider length-matching requirement and wire shape resemblance for high speed signal routes on board. Our results show that we can utilize routing resource very carefully, and can account for the resemblance of nets in the presence of the obstacles. Our approach is workable for board busses as well.
Tsung-Ying Tsai, Ren-Jie Lee, Ching-Yu Chin, Chung-Yi Kuan, Hung-Ming Chen, Yoji Kajitani
DATE5
2011 Fast analog layout prototyping for nanometer design migration
abstract
This paper presents an analog layout migration methodology to quickly provide multiple layouts while keeping similar or better circuit performance. Unlike previous works that often generate a single layout that has exactly the same topology with the original layout, this new migration algorithm is able to provide results with different aspect ratios. First, various placement constraints, including topology, matching, and symmetry, are extracted from the original layout. The extracted constraints are hierarchically stored into a topology slicing tree. Placement is performed from the bottom tree nodes to the root tree node. In each tree node, multiple placements for the subtree are recorded. All possible placements under the constraints are recorded in the root node. This algorithm has been successfully applied to a variable gain amplifier and a folded cascode operational amplifier migrating from UMC 90nm to UMC 65nm. The experimental results validate that our approach can provide reasonable layouts, even a better result almost in no time.
Yi-Peng Weng, Hung-Ming Chen, Tung-Chieh Chen, Po-Cheng Pan, Wei-Zen Chen
ICCAD2
2011 Aseismic ability estimation of school building using predictive data mining models
Wei-Ko Kao, Hung-Ming Chen, Jui-Sheng Chou
Expert Syst. Appl.2
2011 Efficient Package Pin-Out Planning With System Interconnects Optimization for Package-Board Codesign
abstract
In conventional package design, engineers designate the ball grid array (BGA) pin-out manually, this always postpones the time-to-market (TTM) of products due to the turn-around between package and design houses. Recent papers propose a method of automatically generating the pin-out and taking signal integrity (SI), power delivery integrity (PI), and routability (RA) into account simultaneously by pin-block design and floorplanning, thus dramatically speeding up the developing time. However, this approach ignores the considerations of shorter path length and equilength/length matching in routing printed circuit board (PCB) trace and pin-out assignment for high-speed interface IP designs, such as USB and PCI Express. Since these features are the most important performance metrics during chip-package-board codesign, in this paper we propose the ideas to optimize the system interconnects during package pin-out design. These ideas keep the same minimized package size as aforementioned recent work and ensure that SI, PI, and RA can still be considered with significant reduction in design cost. It is achieved by relaxing the restriction of pin-block side and order on the package, usually specified by package designers. The experimental results on industrial chipset design cases show that the average improvement of our pin-block planner is over 40% when comparing the design cost with the previous work, among which we have one case accommodated over a thousand pins. Our ideas also work for any kind of pin-block or pin-group configurations.
Ren-Jie Lee, Hung-Ming Chen
IEEE Trans. Very Large Scale Integr. Syst.2
2010 Technology mapping with crosstalk noise avoidance
abstract
In today's VLSI designs, crosstalk effects causing chips to fail or suffer from low yields have become one of the very essential design issues. In this paper, we attempt to reduce crosstalk noise in logic and physical synthesis stage, which is usually done in post-layout stage. We propose a technology mapping method that can reduce the crosstalk noise while meeting delay constraints. The algorithm employing a dynamic programming framework in the matching phase determines the routing of fanin nets for all the matches to estimate the track utilization in probability. These routings are stored as virtual routing maps to compute the crosstalk noise during the covering phase, which will select the crosstalk-minimal solutions satisfying the delay constraints rather than the delay-minimal ones. This problem is different from wire congestion-driven technology mapping and our experimental results are encouraging. We experiment on the benchmark circuits in 90nm process, the results show that, with 7% of area increase, our proposed approach is effective to improve the crosstalk by 28% on average, as compared to the conventional delay- and/or congestion-driven technology mapping. The overall result is better than the efforts done in post-layout stage, and has been validated by modern commercial EDA tools. In addition, this proposed approach can be applied in local technology remapping at post-placement/post-routing and ECO stages as well.
Fang-Yu Fan, Hung-Ming Chen, I-Min Liu
ASP-DAC2
2010 On Reducing Test Power and Test Volume by Selective Pattern Compression Schemes
abstract
In modern chip designs, test strategies are becoming one of the most important issues due to the increase of the test cost, among them we focus on the large test power dissipation and large test data volume. In this paper, we develop a methodology to suppress the test power to avoid chip failures caused by large test power, and our methodology is also effective in reducing the test data volume and shift-in power. The proposed schemes and techniques are based on the selective test pattern compression, they can reduce considerable shift-in power by skipping the switching signal passing through long scan chains. The experimental results with ISCAS89 circuits demonstrate that our methodology can achieve significant improvement in the reduction of shift-in power and test data volume. Our approach also supports multiple scan chains.
Chia-Yi Lin, Hsiu-Chuan Lin, Hung-Ming Chen
IEEE Trans. Very Large Scale Integr. Syst.3
2009 Package routability- and IR-drop-aware finger/pad assignment in chip-package co-design
abstract
Due to increasing complexity of design interactions between the chip, package and PCB, it is essential to consider them at the same time. Specifically the finger/pad locations affect the performance of the chip and the package significantly. In this paper, we have developed techniques in chip-package codesign to decide the locations of fingers/pads for package routability and signal integrity concerns in chip core design. Our finger/pad assignment is a two-step method: first we optimize the wire congestion problem in package routing, and then we try to minimize the IR-drop violation with finger/pad solution refinement. The experimental results are encouraging. Compared with the randomly optimized methods, our approaches reduce in average 42% and 68% of the maximum density in package and 10.61% of IR-drop for test circuits.
Chao-Hung Lu, Hung-Ming Chen, Chien-Nan Jimmy Liu, Wen-Yu Shih
DATE2
2009 A stochastic-based efficient critical area extractor on OpenAccess platform
abstract
Due to inefficient calculation in current critical area analyzer, in this work we present a method of extracting critical area for short faults from the mask layout of an integrated circuit. The method is based on the concept of sampling framework and the geometry computation of critical area.
Bo-Zhou Chen, Hung-Ming Chen, Li-Da Huang, Po-Cheng Pan
ACM Great Lakes Symposium on VLSI2
2009 Performance-constrained voltage assignment in multiple supply voltage SoC floorplanning
abstract
Using voltage island methodology to reduce power consumption for System-on-a-Chip (SoC) designs has become more and more popular recently. Currently this approach has been considered either in system-level architecture or postplacement stage. Since hierarchical design and reusable intellectual property (IP) are widely used, it is necessary to optimize floorplanning/placement methodology considering voltage islands generation to solve power and critical path delay problems. In this article, we propose a floorplanning methodology considering voltage islands generation and performance constraints. Our method is flexible and can be extended to hierarchical design. The experimental results on some MCNC benchmarks show that our method is effective in meeting performance constraints and can simultaneously consider the tradeoff between power routing cost and total power dissipation.
Meng-Chen Wu, Ming-Ching Lu, Hung-Ming Chen, Jing-Yang Jou
ACM Trans. Design Autom. Electr. Syst.3
2009 Fast Flip-Chip Pin-Out Designation Respin for Package-Board Codesign
abstract
Deep submicrometer effects drive the complication in designing chips, as well as in package designs and communications between package and board. As a result, the iterative interface design has been a time-consuming process. This paper proposes a novel and efficient approach to designating pin-out, which is a package ball chart describing pin locations for flip-chip BGA package when designing chipsets. The proposed approach can not only automate the assignment of more than 200 input/output (I/O) pins on package, but also precisely evaluate package size which accommodates all pins with almost no void pin positions, as good as the one from manual design. Furthermore, the practical experience and techniques in designing such interface has been accounted for, including signal integrity, power delivery and routability. This efficient pin-out designation and package size estimation by pin-block design and floorplanning provides much faster turn around time, thus enormous improvement in meeting design schedule. Our pin-block design contains two major parts. First, we have pin-block construction to locate signal pins within a block along the specific patterns. Six pin patterns are proposed as templates which are automatically generated according to the user-defined constraints. Second, we have pin-blocks grouping to group all pin-blocks into package boundaries. Two alternative pin-blocks grouping strategies are provided for various applications such as chipset and field-programmable gate array (FPGA). The results on two real cases show that our methodology is effective in achieving almost the same dimensions in package size, compared with manual design in weeks, while simultaneously considering critical issues and package size migration in package-board codesign.
Ren-Jie Lee, Hung-Ming Chen
IEEE Trans. Very Large Scale Integr. Syst.2
2008 Efficient and optimal post-layout double-cut via insertion by network relaxation and min-cost maximum flow
abstract
As VLSI design complexity is continuously increasing, the yield loss due to via failure becomes more significant. Adding a redundant via adjacent to each single via is a well-known and highly recommended method to reduce yield loss due to via failure. In this paper, we develop a network-flow-based algorithm in post-layout stage for the redundant via insertion problem. With our novel and efficient approach, we can obtain optimal redundant via insertion solution in improving the manufacturing yield, with minimal fixup if necessary. Moreover, our approach is parallel-processing-friendly and effective in ECO incremental solution due to the nature of network-flow models.
Lun-Chun Wei, Hung-Ming Chen, Li-Da Huang, Sarah Songjie Xu
ACM Great Lakes Symposium on VLSI2
2008 Blockage and voltage island-aware dual-vdd buffered tree construction under fixed buffer locations
abstract
Due to the need of low power methodology in VLSI and SoC designs, voltage island architecture is attracting attentions in design community. However, the corresponding EDA tools development for voltage-island-aware buffered routing is still very few. Recent related studies focused on applying dual Vdd buffers in routing tree construction, however it cannot be applied on a design using voltage island architecture due to the restriction on the ordering of low and high Vdd buffers and the lack of level converter consideration. This paper presents approaches to solving the buffer insertion and level converter assignment problem in the presence of voltage island in a low-power design, especially under the fixed buffer locations. We have implemented and modified one state-of-the-art graph-based approach for this specific routing problem and applied our efficient heuristics (one of themis based on the selection of Steiner points) to further improve the performance, considering the assignment of buffers and level converters/shifters simultaneously. The experimental results show that we can obtain massive speedup over modified prior approach, and even with lower power and delay. Furthermore, as the number of sinks increases, our approach can effectively find feasible solutions, while modified prior approach cannot find solutions within a reasonable runtime
Bruce Tseng, Hung-Ming Chen
ISPD2
2008 Effective decap insertion in area-array SoC floorplan design
abstract
As VLSI technology enters the nanometer era, supply voltages continue to drop due to the reduction of power dissipation, but it makes power integrity problems even worse. Employing decoupling capacitances (decaps) in floorplan stage is a common approach to alleviating supply noise problems. Previous researches overestimate the decap budget and do not fully utilize the empty space of the floorplan. A floorplan usually has a lot of available space that can be used to insert the decap without increasing the floorplan area. Therefore, the goal of this work is to develop a better model to calculate the required decap to solve the power supply noise problem in area-array based designs, and increase the usage of available space in the floorplan to reduce the area overhead caused by decap insertion. The experimental results of this work are encouraging. Compared with previous approaches, our methodology reduces 38% of the decap budget in average for MCNC benchmarks but can still meet the power supply noise requirements. The final floorplan areas with decap are also smaller than the numbers reported in previous works.
Chao-Hung Lu, Hung-Ming Chen, Chien-Nan Jimmy Liu
ACM Trans. Design Autom. Electr. Syst.2
2008 Design Migration From Peripheral ASIC Design to Area-I/O Flip-Chip Design by Chip I/O Planning and Legalization
abstract
Due to higher input/output (I/O) count and power delivery problem in deep submicrometer (DSM) regime, flip-chip technology, especially for area-array architecture, has provided more opportunities for adoption than traditional peripheral bonding design style in high-performance application-specific integrated circuit and microprocessor designs. However, it is hard to tell which technique can provide better design cost edge in usually concerned perspectives. In this paper, we present a methodology to convert a previous peripheral bonding design to an area-I/O flip-chip design. It is based on an I/O buffer modeling and an I/O planning algorithm to legalize I/O buffer blocks with core placement without sacrificing much of the previous optimization in the original core placement. The experimental results have shown that we have achieved better area and I/O wirelength in area-IO flip-chip configuration (especially for pad-limit designs), compared with peripheral bonding configuration in packaging consideration.
Chia-Yi Chang, Hung-Ming Chen
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Fast Flip-Chip Pin-Out Designation Respin by Pin-Block Design and Floorplanning for Package-Board Codesign
abstract
Deep submicron effects drive the complication in designing chips, as well as in package designs and communications between package and board. As a result, the iterative interface design has been a time-consuming process. This paper proposes a novel and efficient approach to designating pin-out for flip-chip BGA package when designing chipsets. The proposed approach can not only automate the assignment of more than 200 I/O pins on package, but also precisely evaluate package size which accommodates all pins with almost no void pin positions, as good as the one from manual design. Furthermore, the practical experience and techniques in designing such interface has been accounted for, including signal integrity, power delivery and routability. This efficient pin-out designation and package size estimation by pin-block design and floorplanning provides much faster turn around time, thus enormous improvement in meeting design schedule. The results on two real cases show that our methodology is effective in achieving almost the same dimensions in package size, compared with manual design in weeks, while simultaneously considering critical issues in package-board codesign. To the best of our knowledge, this is the first attempt in solving flip-chip pin-out placement problem in package-board codesign.
Ren-Jie Lee, Ming-Fang Lai, Hung-Ming Chen
ASP-DAC3
2007 On Increasing Signal Integrity with Minimal Decap Insertion in Area-Array SoC Floorplan Design
abstract
With technology further scaling into deep submicron era, power supply noise become an important problem. Power supply noise problem is getting worse due to serious IR-drop and simultaneous switching noise, and decoupling capacitance (decap) insertion is commonly applied to alleviate the noise. There exist some approaches to addressing this issue, but they suffer either from over-design problem or late decap insertion during design stage. In this paper, we propose a methodology to insert decap in a more efficient and effective way during early design stage in area-array designs. The experimental results are encouraging. Compared with other approaches in (Zhao et al., 2002) and (Yan et al., 2005), we have inserted enough decap to meet supply noise constraint while others employ more area.
Chao-Hung Lu, Hung-Ming Chen, Chien-Nan Jimmy Liu
ASP-DAC2
2007 A selective pattern-compression scheme for power and test-data reduction
abstract
This paper proposes a selective pattern-compression scheme to minimize both test power and test data volume during scan-based testing. The proposed scheme will selectively supply the test patterns either through the compressed scan chain whose scanned values will be decoded to the original scan cells, or directly through the original scan chain using minimum transition filling method. Due to shorter length of a compressed scan chain, the potential switching activities and the required storage bits can be both reduced. Furthermore, the proposed scheme also supports multiple scan chains. The experimental results demonstrate that, with few hardware overhead, the proposed scheme can achieve significant improvement in shift-in power reduction and large amount of test data volume reduction.
Chia-Yi Lin, Hung-Ming Chen
ICCAD2
2006 Scoring Method for Tumor Prediction from Microarray Data Using an Evolutionary Fuzzy Classifier
Shinn-Ying Ho, Chih-Hung Hsieh, Kuan-Wei Chen, Hui-Ling Huang, Hung-Ming Chen, Shinn-Jang Ho
PAKDD5
2006 Markov model fuzzy-reasoning based algorithm for fast block motion estimation
Po-Hung Chen, Hung-Ming Chen, Kuo-Jui Hung, Wen-Hsien Fang, Mon-Chau Shie, Feipei Lai
J. Vis. Commun. Image Represent.2
2006 I/O Clustering in Design Cost and Performance Optimization for Flip-Chip Design
abstract
Input-output (I/O) placement has always been a concern in modern integrated circuit design. Due to flip-chip technology, I/O can be placed throughout the whole chip without long wires from the periphery of the chip. However, because of I/O placement constraints in design cost (DC) and performance, I/O buffer planning becomes a pressing problem. During the early stages of circuits and package co-design, I/O layout should be evaluated to optimize DC and to avoid product failures. The objective of this brief is to improve the existing/initial standard cell placement by I/O clustering, considering DC reduction and signal integrity preservation. The authors formulate it as a minimum cost flow problem that minimizes alphaW+betaD, where W is the I/O wirelength of the placement and D is the total voltage drop in the power network and, at the same time, reduces the number of I/O buffer blocks. The experimental results on some Microelectronics Center of North Carolina benchmarks show that the author's method averagely achieves better timing performance and over 32% DC reduction when compared with a conventional rule-of-thumb design that is popularly used by circuit designers
Hung-Ming Chen, I-Min Liu, Martin D. F. Wong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2005 Efficient gene selection for classification of microarray data
abstract
Microarray is a useful technique for measuring expression data of thousands of genes simultaneously. One of challenges in classification of microarray data is to select a minimal number of relevant genes which can maximize classification accuracy. Many gene selection methods as well as their corresponding classifiers have been proposed. One of existing analysis methods is the hybrid approach based on genetic algorithm and maximum likelihood classification (GA/MLHD). In this paper, an intelligent genetic algorithm (IGA) using control genes and an improved fitness function is proposed to determine the minimal number of relevant genes and identify these genes, while maximizing classification accuracy simultaneously. The experimental results show that our approach is superior to the existing method GA/MLHD in terms of the number of selected genes, classification accuracy, and robustness of selected genes and accuracy, especially for the datasets which have numerous categories and a large number of testing genes inside.
Shinn-Ying Ho, Chong-Cheng Lee, Hung-Ming Chen, Hui-Ling Huang
Congress on Evolutionary Computation3
2005 Flexible protein-ligand docking using particle swarm optimization
abstract
Many protein-ligand docking problems attempt to predict the bound conformations of two interacting molecules. Consequently, the docking problem requires a powerful search technique to explore the translations, orientations, and each torsion until an ideal site has been found. Therefore, protein-ligand docking can be formulated as a parameter optimization problem. However, highly flexible ligands have a lot of torsions. Therefore, the optimization problem of highly flexible docking would become more difficult due to the increment of parameter number and interactions among these parameters. We proposed a novel method SODOCK based on particle swarm optimization (PSO) for solving flexible protein-ligand docking problems. PSO has significant effect on the optimization of parameters with strong interactions. A commonly used efficient local search is incorporated into SODOCK to improve the efficiency and robustness of PSO. SODOCK is efficient for both types of ligands with small and large numbers of torsions. It is shown by computer simulation that SODOCK performs well in obtaining accurate conformations, compared with some of state-of-the-art methods. Moreover, it is also shown that SODOCK is superior to AutoDock using the same energy function in AutoDock 3.05 in terms of convergence speed, robustness, and docking energy, especially for highly flexible docking problems
Bo-Fu Liu, Hung-Ming Chen, Hui-Ling Huang, Shiow-Fen Hwang, Shinn-Ying Ho
Congress on Evolutionary Computation2
2005 MeSwarm: memetic particle swarm optimization
abstract
In this paper, a novel variant of particle swarm optimization (PSO), named memetic particle swarm optimization algorithm (MeSwarm), is proposed for tackling the overshooting problem in the motion behavior of PSO. The overshooting problem is a phenomenon in PSO due to the velocity update mechanism of PSO. While the overshooting problem occurs, particles may be led to wrong or opposite directions against the direction to the global optimum. As a result, MeSwarm integrates the standard PSO with the Solis and Wets local search strategy to avoid the overshooting problem and that is based on the recent probability of success to efficiently generate a new candidate solution around the current particle. Thus, six test functions and a real-world optimization problem, the flexible protein-ligand docking problem are used to validate the performance of MeSwarm. The experimental results indicate that MeSwarm outperforms the standard PSO and several evolutionary algorithms in terms of solution quality.
Bo-Fu Liu, Hung-Ming Chen, Jian-Hung Chen, Shiow-Fen Hwang, Shinn-Ying Ho
GECCO2
2005 Design of nearest neighbor classifiers: multi-objective approach
Jian-Hung Chen, Hung-Ming Chen, Shinn-Ying Ho
Int. J. Approx. Reason.2
2005 Simultaneous power supply planning and noise avoidance in floorplan design
abstract
With today's advanced integrated circuit manufacturing technology in deep submicron (DSM) environment, we can integrate entire electronic systems on a single system on a chip. However, without careful power supply planning in layout, the design of chips will suffer from local hot spots, insufficient power supply, and signal integrity problems. Postfloorplanning or postroute methodologies in solving power delivery and signal integrity problems have been applied but they will cause a long turnaround time, which adds costly delays to time-to-market. In this paper, we study the problem of simultaneous power supply planning and noise avoidance as early as in the floorplanning stage. We show that the problem of simultaneous power supply planning and noise avoidance can be formulated as a constrained maximum flow problem and present an efficient yet effective heuristic to handle the problem. Experimental results are encouraging. With a slight increase of total wirelength, we achieve almost no static IR (voltage)-drop requirement violation in meeting the current and power demand requirement imposed by the circuit blocks compared with a traditional floorplanner and 45.7% of improvement on a /spl Delta/I noise constraint violation compared with the approach that only considers power supply planning.
Hung-Ming Chen, Li-Da Huang, I-Min Liu, Martin D. F. Wong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2004 I/O Clustering in Design Cost and Performance Optimization for Flip-Chip Design
abstract
I/O placement has always been a concern in modern IC design. Due to flip-chip technology, I/O can be placed throughout the whole chip without long wires from the periphery of the chip. However, because of I/O placement constraints in design cost and performance, I/O buffer planning becomes a pressing problem. During the early stages of circuits and packaging co-design, I/O layout should be evaluated to optimize design cost and to avoid product failures. In this paper, our objective is to better an existing/initial standard cell placement by I/O clustering, considering design cost reduction and signal integrity preservation. We formulate it as a minimum cost flow problem minimizing /spl alpha/W+/spl beta/D, where W is the I/O wirelength of the placement and D is the total voltage drop in the power network. The experimental results on some MCNC benchmarks show that our method achieves better timing performance and averagely over 30% design cost reduction when compared with the conventional design rule of thumb popularly used by circuit designers.
Hung-Ming Chen, I-Min Liu, Martin D. F. Wong, Muzhou Shao, Li-Da Huang
ICCD1
2004 Design of Nearest Neighbor Classifiers Using an Intelligent Multi-objective Evolutionary Algorithm
Jian-Hung Chen, Hung-Ming Chen, Shinn-Ying Ho
PRICAI2
2004 Design of accurate classifiers with a compact fuzzy-rule base using an evolutionary scatter partition of feature space
abstract
An evolutionary approach to designing accurate classifiers with a compact fuzzy-rule base using a scatter partition of feature space is proposed, in which all the elements of the fuzzy classifier design problem have been moved in parameters of a complex optimization problem. An intelligent genetic algorithm (IGA) is used to effectively solve the design problem of fuzzy classifiers with many tuning parameters. The merits of the proposed method are threefold: 1) the proposed method has high search ability to efficiently find fuzzy rule-based systems with high fitness values, 2) obtained fuzzy rules have high interpretability, and 3) obtained compact classifiers have high classification accuracy on unseen test patterns. The sensitivity of control parameters of the proposed method is empirically analyzed to show the robustness of the IGA-based method. The performance comparison and statistical analysis of experimental results using ten-fold cross validation show that the IGA-based method without heuristics is efficient in designing accurate and compact fuzzy classifiers using 11 well-known data sets with numerical attribute values.
Shinn-Ying Ho, Hung-Ming Chen, Shinn-Jang Ho, Tai-Kang Chen
IEEE Trans. Syst. Man Cybern. Part B2
2003 Floorplanning with power supply noise avoidance
abstract
Abstract — With today’s advanced integrated circuits (ICs) manufacturing technology in deep submicron (DSM) environ-ment, we can integrate entire electronic systems on a single chip (SoC). However, without careful power supply planning in lay-out, the design of chips will suffer from mostly signal integrity problems including IR-drop, I noise, and IC reliability. Post-route methodologies in solving signal integrity problem have been applied but they will cause a long turn-around time, which adds costly delays to time-to-market. In this paper, we study the prob-lem of power supply noise avoidance as early as in floorplanning stage. We show that the noise avoidance in power supply planning problem can be formulated as a constrained maximum flow prob-lem and present an efficient yet effective heuristic to handle the problem. Experimental results are encouraging. With slight in-crease of total wirelength, we achieve almost no IR-drop require-ment violation and 46.6 % of improvement on I noise constraint violation compared with a previous approach. I.
Hung-Ming Chen, Li-Da Huang, I-Min Liu, Minghorng Lai, Martin D. F. Wong
ASP-DAC1
2003 Global Wire Bus Configuration with Minimum Delay Uncertainty
Li-Da Huang, Hung-Ming Chen, Martin D. F. Wong
DATE2
2001 Integrated power supply planning and floorplanning
abstract
One of the most challenging issues in today's high-performance VLSI design is to ensure high-quality power supply to each individual circuit blocks. Reduced power supply voltage can result in slower cell switching, or even circuit failure. Nevertheless, most floorplanning methodologies have ignored power supply considerations. Thus, the resulting floorplan may suffer from local hot spots and insufficient power supply for certain circuit blocks. In this paper, we present an optimal power supply planning algorithm based on network flow to shorten the current paths from power bumps to local power supply wirings. We have incorporated our algorithm into a floorplanning algorithm for integrated floorplanning and power supply planning. Experimental results are encouraging.
I-Min Liu, Hung-Ming Chen, Tan-Li Chou, Adnan Aziz, Martin D. F. Wong
ASP-DAC2
2001 Faster and more accurate wiring evaluation in interconnect-centric floorplanning
abstract
#%$& (' ) * + , . /'0* 1 2 * 34 657 8* 2 8* 9 /':; : 6 * ?6@8AB *1 1 0* %* C= 8 0 ?D':1 ? E3F * +G'H ?D'0* IJ ?6@ K'0? ?L > M 2 4 0NO ? 8 (P> Q -A : A )7R-SJ$& *F+='T'Q?U'V : > M 2 4 0N * A : AXW-S0YM$(A2!> , [Z \* :MC= 8 '0?D:6* + F M ?D'0* / ]'0 ?D':^3F+ (+] E* % Z 'K' ?D'0 :H > `_ 2 [ . 0NTC 8 ?U'0 )a* + K [ /':% > M 2 -NT * M+ 'M K' 8* * _b 8* 9C 8 ?U'0 D :c *(''-? ?6@]IJ @d Z>_ 2 IJ Ae#1 0 VIJ /)F* + 1 D 9'0?D \*9 G \@ \* K'V* D H3f' @;* ? IJ .* + ` : \* d ? g (+c K'0: * M -N4 > `_ 2 [ h 0NO [* T ^'M: IJ 9C= 8 ?D'LAi H* + F ='0 2 /) 34 E 8*T' ?D E@8 [*h [57 * IJ E D /'.* :6 ('K'0* 4 0I 8*r Q 8* AOx 0 '.y y0_b ?D 8 (Pcz/R0Y,_b [*F ? ) 3X , / / K 8* hNu { 0I W0y`+ 4* `? F* +='0 ^R-SM D > * A 1. INTRODUCTION n| * +H},qr! a* -?D -:-@H 8* :Q* + E 1 c ! #%$ [ (' )r I> E'V . /'0? / c /3F * 9 K'-? ? E ~ = % ?D'/ c'V* '0 ^ IJ h ':9 /Z * @8Af#1 '83F+ ? )73F 6* + * + /'0 0No h ~ ) fNu [* D f'V h 8* :0 ('0* / K >* T D tAih? ? * + E : [ G},qr! , +>@ /'0?4 [_ : %v6z/S>) „0wAak /34 IJ /)J3F 6* +`* & 0 X+> / a -N7* + 'i -N 8* [*f H \*('0 = 'V ( _b ? ?2 0 4Np ?D?6_b \*  : H'0Nu* [ F 6 _ * L ='V * * : ) 8* *r ?D':F L* _b : A # '8@cC= 8 '-? : 0 6* + )i ? = :% ? D D : '= G _ ? : )7+ '/IJ E 2 2 / ^ ='\*k /' Mv6z-z )Oz  )tW>)2‘>)7’>) RVw)> *a* + @` D ` 0*&*('0PJ F >* [ *& ?D'0 :E D 8* Q'>* Z *9v W/wA € ;v WVw)&ƒX+ d *M'-? A1 8* :0 ('0* C= 8 0 ?D':%'0 = “ D >*M 8* [*. ?U'0 D : Ac!> D IJ [ @d3F 6 :% I-'-? ='V_ * 9 \* K X'Q ?D [_ :[* @`: ? ='-?2 * /)1v W/w7 /'9 ?6@ 8 f > ('V* _b ~ / * ? \* , A : ATz S S-SM [* [$a *X3F ? ?7 -* 2 “ D >*Q* -?DI 9 ?D 6* +]IJ [ @c?D'V : K * ? \* 1 -A : A z/S0Y” [* [$(A,x * + 0 )ˆv WVwa 'K+ \* M:-?D ='0?o * [ * K /m8 8* D'-? ?6@^ * M:-? ='0?O >* [ * )7* + Q 2 [ Np K'0 -NO3F+ [+1 2 f 1* + , 0 ( : -NO * f* . 2 , * / tA c [Z \* :]C 8 ?D'0 :;'0? : 6* + . ?D'V* / Š'_ '-? :h3F+ [+ / r* T Z '&'F?D'0 :4 > . 2 [ r 0N C 8 ?D'0 ) * + f /':E . 2 o -N7 [* i+ 'i K' 4 >* [ * _b >* C= 8 0 ?D': *(''-? ?6@ IJ @. Z 2 IJ Ao € `* + i ='0 2 /) 34 c >*1'‰ D ? %@J * 57 * IJ c D /'G* ˆ :6 c 8* Š'Œ :-? ˆ */A•…X Nu ‰::|* + G * ) . ?6* _* '-?o * E'0 M * 9*€3X -_* ='0? [* A …X = / _ :M+8@ 2 [ :('+ _* -_b:0 ('+%* ('\Np 0 K'0* . / d* % I K* + K \* ('8*.Np 0 _ ? 6*Q D ‰'^ ?D 8 (P2A l4+ E *, / * %* (+ Um8 . ,IJ @% [57 * IJ A.! 2 M34 '0 h: IJ H' ? ‹3F 6* +1R-S ? 8 [P> f'0 = 9W-S0YŽ [* AiTNu* [ f * / * t)834 , 'K+='/IJ h'0*f \*,zVW-R-SQ * h p3F+ [+9 X* + h /'3F+ `34 F+ 'VI f * i 2 *€3X K'-? ? ='6 i -Nt ? 8 (P> [$(Aal4+ Q'H :6 * )-Np 0 F', _ ? •3F * +Œy y ? 8 [P> '0 = ŠzVRVY† * )&* + HC 8 ?U'0 D :G'0?D:-_ 6* +  D cv WVwr* 8 -PH 0 k* +='0 ^W0y.+ f* M 1 H* + , 0 D:='-? * ? \*/)& ** 8 Pc? M* + '‰R0S% > * '0Nu* [ ` [*M / * LA #1 '83F+ ? )k3X c 8*H'] %'('V* c:-?D ='0?, * [ HNu 3F 6 :^ I '0?D '0* d D * + E ='0 2 /A nG . qL'-:('0 :U'0 G ?D'0Z>_ 'V* .* (+ Dm> &* h \@ \* K'0* /'-? ?6@Q * X: ? '-? D 8* * ) * @ :,* E D ~ 4* + f K'0Z . ŽI> D -?D'0* 9'-: '\*&* + 4 * _ :k Ail4+ D q'%: IJ ‰C 8 ?U'0 LA‰l4+ H 7 'V* H 0N,qL'-:('0 :U'0 b2 b5
Hung-Ming Chen, Martin D. F. Wong, Wai-Kei Mak, Hannah Honghua Yang
ACM Great Lakes Symposium on VLSI1
1999 Integrated floorplanning and interconnect planning
abstract
VLSI fabrication has entered the deep sub-micron era and communication between different components has significantly increased. Interconnect delay has become the dominant factor in total circuit delay. As a result, it is necessary to start interconnect planning as early as possible. We propose a method to combine interconnect planning with floorplanning. Our approach is based on the Wong-Liu (1986) floorplaning algorithm. When the positions, orientations, and shapes of the cells are decided, the pin positions and routing of the interconnects are decided as well. We use a multi-stage simulated annealing approach in which different interconnect planning methods are used in different ranges of temperature to reduce running time. A temperature adjustment scheme is designed to give smooth transitions between different stages of simulated annealing. Experimental results show that our approach performs well.
Hung-Ming Chen, Hai Zhou 0001, Evangeline F. Y. Young, Martin D. F. Wong, Hannah Honghua Yang, Naveed A. Sherwani
ICCAD1