EDBT 2026 Demo / reviewers in the wild / expert
Charlie Chung-Ping Chen
dblp:c/CharlieChungPingChen · also Chung-Ping Chen
· DBLP profile ↗
86ranked-venue papers
10as first author
4since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 82 · 9 first-author · 3 since 2021Software engineering, systems software and programming languages · 11 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Transformer Based Real Time Musculoskeletal Anatomical Structure Detection in Clinical UseabstractMedical ultrasound imaging is known as a non-invasive and radiation-free real-time imaging method, making it an indispensable technique for clinical use, especially in musculoskeletal medicine. However, due to the complexity of the body's internal anatomical structures, it is challenging for physicians to rapidly and accurately interpret images to identify each structure for further diagnosis and treatment planning. To address these challenges, we developed a real-time labeling system of anatomical structures for musculoskeletal ultrasound using a transformer-based model, RT-DETR. We selected the 15 most crucial anatomical structures in the human forearm and shoulder as our identification targets. We achieved an overall precision of 0.834, recall of 0.822, mAP50 of 0.855, mAP50:95 of 0.526, and an F1-score of 0.82. Moreover, the real-time annotation results also show high precision in annotating each structure's classes and boundaries, indicating that our model is well-suited for the task and has the potential to be applied in clinical settings. Jyun-Ping Kao, Hao-Yu Hung, Ping-Xuan Chen, Charlie Chung-Ping Chen, Wen-Shiang Chen |
BIBE | 4 |
| 2023 | A Ripple-Based Constant On-Time Controlled DC-DC Buck Converter with Inductor Current Sensing TechniqueabstractThis letter implements a ripple-based constant on-time (RBCOT) buck converter with a fast transient response fabricated in the TSMC$0.18 \mu \mathrm{m}$CMOS process. The steady-state measurement shows that this chip can regulate output voltage from 0.9V to 1.8V while the input voltage is set at 3.3V and the output load current varies from 0.1A to 1A. The load transient response shows that when the output voltage is 0.9V, the undershoot voltage is 78m V and the overshoot voltage is 126m V. The settling time is$2.8\mu \mathrm{s}$for a step-up load and$2.5\mu \mathrm{s}$for a step-down load. The chip area is 0.922 mm2, The maximum efficiency is 92.32%. This work improves the traditional inductor current ramp compensation technique. Here, we utilize an accurate transconductance amplifier to amplify the inductor current, increasing system stability and efficiency. By adopting negative inductor current feedback to improve the transient response. Through the V2controlled dual loop structures to eliminate the output dc voltage offset issues. Sheng-Jen Cheng, Chieh-Ju Tsai, Sheng-Yu Wang, Charlie Chung-Ping Chen |
ISCAS | 5 |
| 2022 | Intelligent Design Automation for Heterogeneous IntegrationabstractAs the design complexity grows dramatically in modern circuit designs, 2.5D/3D heterogeneous integration (HI) becomes effective for system performance, power, and cost optimization, providing promising solutions to the increasing cost of more-Moore scaling. In this talk, we investigate the chip, package, and board co-design methodology with advanced packages and optical communication considering essential issues on physical design, electrical, thermal, and mechanical effects, timing, and testing, and suggest future research opportunities. Layout: A robust and vertically integrated physical design flow for HI design is needed. We address chip-, package-, and board-level component planning, package-level RDL routing, board-level routing, optical routing, and placement and routing considering warpage and thermal effects. Timing: New chip-level and cross-chip timing analysis techniques are desired. We address timing propagation under current source delay model (CSM), timing analysis and optimization for optical-electrical routing, multi-corner multi-mode analysis for HI, hierarchical MCMM analysis. Testing: The scope covers functional-like test generation, System-in-Package (SiP) online testing, photonic integrated circuits (PIC) testing and design-for-test (DfT), etc. Integration: We shall address chip, package, and board co-design considering multi-domain physics, including physical, electrical, thermal, mechanical, and optical effects and optimization. Iris Hui-Ru Jiang, Yao-Wen Chang, Jiun-Lang Huang, Charlie Chung-Ping Chen |
ISPD | 4 |
| 2021 | Power Reduction of a Set-Associative Instruction Cache Using a Dynamic Early Tag LookupabstractAn energy-efficient instruction cache lookup technique with low area overheads is proposed. The key concept of this Dynamic Early Tag Lookup (DETL) method is to exploit the presence of instruction fetch-bubble cycles. In a fetch-bubble cycle, the index of the matching cache set can be determined earlier. Hence, the dynamic energy for parallel memory accesses to irrelevant cache banks can be saved. We implemented the proposed DETL algorithm on a 4-way set-associative instruction cache in a RISC-V micro-architecture, and tested its performance using the SPEC CPU2006 benchmark suite. The experiment results showed a 19.38% dynamic power reduction with an area overhead smaller than 0.1 %. Chun-Chang Yu, Yu Hen Hu, Yi-Chang Lu, Charlie Chung-Ping Chen |
DATE | 4 |
| 2020 | Intelligent Design Automation for 2.5/3D Heterogeneous SoC IntegrationabstractAs the design complexity grows dramatically in modern circuit designs, 2.5D/3D chip/package/board integration has become a key to beat process limitation for optimizing system performance and power consumption. Among the explored technologies, the wafer-level integrated fan-out (InFO) package-on-package (PoP) has been adopted by major companies such as TSMC to achieve high-density, high-performance, low-cost packaging solutions. To achieve a high-quality 2.5D/3D heterogeneous integration system, we shall study the chip, package, and board codesign methodology with advanced packages and explore key techniques to handle the emerging challenges in physical design, timing, electrical effects, and testing.1 Iris Hui-Ru Jiang, Yao-Wen Chang, Jiun-Lang Huang, Charlie Chung-Ping Chen |
ICCAD | 4 |
| 2017 | An efficient DFT-based algoritiim for the charger noise problem in capacitive touch applicationsabstractOwing to charger noise, inaccurate touch points and fake touch points may happen so that a device has a wrong behavior. The intensity of charger noise could be much larger than that of original touch signals. Besides, the frequency of charger noise varies for each different charger. Therefore, industry identifies charger noise as the most difficult problem in capacitive touch applications. The solution to the problem has become a key in the mobile market. In this paper, we prove that the combination of frequency hopping and repeated integration is an effective method to handle the problem. In addition, we propose an efficient discrete Fourier transform (DFT)-based algorithm to select a good sampling frequency. Moreover, we show an efficient hardware-software co-design adopting our method in Touch IC. Experimental results show that our method can increase SNR by 37 dB and find a good sampling frequency fast and dynamically. Shih-Lun Huang, Sheng-Yi Hung, Charlie Chung-Ping Chen |
ISCAS | 3 |
| 2016 | Lossless compression algorithm based on dictionary coding for multiple e-beam direct write system
Pei-Chun Lin, Yu-Hsuan Pai, Yu-Hsiang Chiu, Shao-Yuan Fang, Charlie Chung-Ping Chen |
DATE | 5 |
| 2015 | Clustering-based multi-touch algorithm framework for the tracking problem with a large number of points
Shih-Lun Huang, Sheng-Yi Hung, Charlie Chung-Ping Chen |
DATE | 3 |
| 2015 | A 8.1/5.4/2.7/1.62 Gb/s receiver for DisplayPort Version 1.3 with automatic bit-rate tracking schemeabstractIn this paper, a wide range and low power multi-rate receiver for DisplayPort Version 1.3 is proposed. In order to extend the bandwidth, a high speed AC coupled interconnect receiver comprising output compensated negative impedance and positive feedback techniques is introduced. Furthermore, the automatic bit-rate tracking scheme is used for clock and data recovery (CDR) to achieve wide data rate range. Besides, this wide range CDR is realized by omitting the power-hungry divider. Thus, the required area and the corresponding power consumption can be substantially reduced. Designed and fabricated in 90nm CMOS technology, this test chip occupies 0.23 mm2and consumes 90 mW. The measured root-mean-square jitter is 5.52/3.15/2.96/3.6 psrmswith the data rates of 8.1/5.4/2.7/1.62 Gb/s, respectively. The bit error rate (BER) for all data rate is less than 10-12for 27-1 pseudo random binary sequences (PRBS). Ai Chien, Shuo-Hong Hung, Kuan-I Wu, Chang-Yi Liu, Min-Han Hsieh, Charlie Chung-Ping Chen |
ISCAS | 6 |
| 2015 | A 160MHz-to-2GHz low jitter fast lock all-digital DLL with phase tracking techniqueabstractAn all-digital delay-locked loop (ADDLL) is proposed for wide range, fast lock, low jitter and high process-voltage-temperature (PVT) tolerance. The proposed phase tracking generator (PTG) produces two tracking rising and falling phases in only 2 cycles for fast lock and wide-range. The digital phase interpolator (DPI) and the control block are adopted to calibrate the phase offsets and random jitters while maintaining the closed-loop property that allow for tracking of PVT variations. The wide-range ADDLL operates from 160MHz to 2GHz. The measured peak-to-peak jitters are 6.89ps and 16.67ps at 2GHz and 160MHz. This chip is fabricated in TSMC 90nm CMOS technology with an active area of 0.205mm2. Shuo-Hong Hung, Wei-Hao Kao, Kuan-I Wu, Yi-Wei Huang, Min-Han Hsieh, Charlie Chung-Ping Chen |
ISCAS | 6 |
| 2015 | A fast-settling high linearity auto gain control for broadband OFDM-based PLC systemabstractA fast-settling high-linearity automatic gain control (AGC) for broadband OFDM-based (Orthogonal Frequency Division Multiplexing) powerline communication (PLC) transceiver is presented in this paper. The high peak-to-average power ratio (PAPR) of OFDM signals makes conventional closed-loop AGCs impractical for the stringent settling-time constraints. A novel AGC algorithm is proposed incorporating with a linear power detector and a charge-redistribution converter to determine the gain of the pseudo-exponential programmable gain amplifier (PGA). Fabricated in 90-nm CMOS technology, the PGA has been measured to exhibit a constant 3-dB bandwidth of 350 MHz with dB-linear gain ranging from -20 to 20 dB. The measured settling time of gain acquisition is less than 0.1 μs, which meets the stringent criteria of short settling time of OFDM-based system. The proposed AGC consumes 3.7 mW through a 1.2-V supply, and occupies 0.07 mm2active area. Kuan-I Wu, Szu-Yao Hung, Shuo-Hong Hung, Charlie Chung-Ping Chen |
ISCAS | 4 |
| 2014 | Current-mode adaptively hysteretic control for buck converters with fast transient response and improved output regulationabstractThis paper presents a current-mode adaptively hysteretic control (CMAHC) technique to achieve the fast transient response for DC-DC buck converters. A complementary full range current sensor comprising of chargingpath and discharging-path sensing transistors is proposed to track the inductor current seamlessly. With the proposed current-mode adaptively hysteretic topology, the inductor current is continuously monitored, and the adaptively hysteretic threshold is dynamically adjusted according to the feedback information comes from the output voltage level. Therefore, a fast load-transient response can be achieved. Besides, the output regulation performance is also improved by the proposed dynamic current-scaling circuitry (DCSC). Moreover, the proposed CMAHC topology can be used in a nearly zero RESR design configuration. The prototype fabricated using TSMC 0.25μm CMOS process occupies the area of 1.78mm2 including all bonding pads. Experimental results show that the output voltage ripple is smaller than 30mV over a wide loading current from 0 mA to 500 mA with maximum power conversion efficiency higher than 90%. The recovery time from light to heavy load (100 to 500 mA) is smaller than 5μs. Kuan-I Wu, Shuo-Hong Hung, Shang-Yu Shieh, Bor-Tsang Hwang, Szu-Yao Hung, Charlie Chung-Ping Chen |
ISCAS | 6 |
| 2013 | A 10-bit current-steering DAC for HomePlug AV2 powerline communication system in 90nm CMOSabstractA 10-bit current-steering digital-to-analog converter (DAC) has been proposed. This study is used for the transmitter (Tx) of the powerline communication (PLC) analog-front-end (AFE), and it reaches the standard of the HomePlug AV2. The proposed DAC also uses a technique as digital random-return-to-zero (DRRZ) [1] to achieve high performance in the OFDM communication systems. The test chip was fabricated in TSMC 90 nm CMOS technology and occupied 0.42 mm2for active area. The supplies for the analog and digital circuits are 2.5V and 1.2V. The maximum INL and DNL are 0.8 LSB and 0.3 LSB respectively. The SFDR is up to 40.09 dB for 1.25GS/s of Nyquist-rate sampling. Wei-Sheng Cheng, Min-Han Hsieh, Shuo-Hong Hung, Szu-Yao Hung, Charlie Chung-Ping Chen |
ISCAS | 5 |
| 2013 | A high dynamic range programmable gain amplifier for HomePlug AV powerline communication systemabstractA high dynamic range programmable gain amplifier (PGA) for HomePlug AV powerline communication (PLC) receivers has been designed in a standard 0.18-μm CMOS technology. Linear transconductance amplifiers are employed as gain cells to obtain high dynamic range. A current compensation technique is proposed to achieve fine gain step characteristics. For constant settling time issue in automatic gain control (AGC) loop, a pseudo-exponential approximation is adopted. A 6-bit PGA suited for PLC receiver is realized by cascading PGA gain cell along with fixed gain amplifiers. The proposed PGA has been measured to exhibits a 3-dB bandwidth of more than 78 MHz and a dB-linear gain ranging from -24 to 74 dB with a maximum gain error of 0.24LSB. The core circuit consumes 1.9 mW through a 1.8 V supply, and occupies only 0.03 mm2active area. Szu-Yao Hung, Kai-Hsiang Chan, Charlie Chung-Ping Chen |
ISCAS | 3 |
| 2013 | A 52 dBc MTPR line driver for powerline communication HomePlug AV standard in 0.18-μm CMOS technologyabstractIn this paper, a line driver for HomePlug AV powerline communication system has been described. The proposed line driver includes a damping factor control (DFC) network which suppressed the open loop high frequency peaking effect to improve the stability for the various characteristic impedance of powerline. The line driver was fabricated in TSMC 0.18-μm CMOS technology and occupied 0.195 mm2active area. When operating with a power line and a coupling unit including a 1:4 turn ratio transformer, the line driver achieves 60 MHz bandwidth and 52.14 dBc in-band MTPR, with 5.15 dBm output signal power and 4.44 peak-to-average ratio (PAR), while 2.5 V supply voltage. Pang-Kai Liu, Szu-Yao Hung, Chang-Yi Liu, Min-Han Hsieh, Charlie Chung-Ping Chen |
ISCAS | 5 |
| 2012 | A 2 - 8 GHz multi-phase distributed DLL using phase insertion in 90 nmabstractA 2 GHz to 8 GHz wide-range multi-phase distributed delay-locked loop (DDLL) has been proposed. The architecture achieves wide operating frequency range by adding a digital phase selector into the DDLL [1]. The insertion phases which are generated from digital phase selector with minor phase error could be fine-tuned by the voltage-controlled delay cells in the DDLL independently. The test chip was fabricated in TSMC 90 nm technology and occupies 0.0644 mm2active areas. The maximum phase error is 1.53 ps at 8 GHz, and 1.93 ps at 2 GHz respectively. Min-Han Hsieh, Bing-Feng Lin, Yu-Shun Wang, Hao-Huei Chang, Charlie Chung-Ping Chen |
ISCAS | 5 |
| 2012 | Efficient Thermal Simulation for 3-D IC With Thermal Through-Silicon ViasabstractA novel virtual power source (VPS) method is proposed that can significantly speedup 3-D integrated circuit (IC) thermal simulation with sparely deployed thermal TSVs leveraging Green function-based analytical spectral method. Specifically, it is shown that the impact of TSVs with nonhomogeneous thermal conductivity on thermal simulation may be modeled as VPSs located at mesh grids inside the thermal TSVs. Using this VPS model, the temperature distribution over the entire 3-D IC may be obtained by solving for the thermal distribution of a thermally homogeneous substrate in the presence of these VPSs. As such, the fast spectral method may be applied to accelerate the simulation. Significant (3–100 times) speedup over a baseline finite difference method implementation has been observed in preliminary simulations. Dongkeun Oh, Charlie Chung-Ping Chen, Yu Hen Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2011 | Epileptic Seizure Detection for Multichannel EEG Signals with Support Vector MachinesabstractEpilepsy is a common chronic neurological disorder characterized by recurrent unprovoked seizures. The electroencephalogram (EEG) signals play an important role in the diagnosis of epilepsy. In addition, multi-channel EEG signals have much more discrimination information than a single channel. However, traditional recognition algorithms of EEG signals are lack of multichannel EEG signals. In this paper, we propose a new method of epileptic seizure detection based on multichannel EEG signals. Both unipolar and bipolar EEG signals are considered in our approach. We make use of approximate entropy (ApEn) and statistic values to extract features. Furthermore, we tested the performance of four different Support Vector Machines (SVMs). The results reveal that the grid SVM achieves the highest totally classification accuracy (98.91%). Chia-Ping Shen, Chih-Min Chan, Feng-Sheng Lin, Ming-Jang Chiu, Jeng-Wei Lin, Jui-Hung Kao, Charlie Chung-Ping Chen, Feipei Lai |
BIBE | 7 |
| 2011 | A 1.2V 6.4GHz 181ps 64-bit CD domino adder with DLL measurement techniqueabstractA novel 64-bit hybrid radix-4 sparse-4 tree adder using clock-delayed (CD) footless domino logic is proposed. The adder operates at 6.4GHz with 181ps latency and it consumes 840mW at 1.2V in a standard 90nm CMOS technology. The adder latency is accurately measured by the programmable clock generated from delay-locked loop (DLL). Pseudo-exhaustive testing is applied so that all testable faults in this 64-bit adder are detected by just 23K patterns. This at-speed self testing technique is very useful for speed binning of high performance CPU. Yu-Shun Wang, Min-Han Hsieh, Chia-Ming Liu, Yi-Chi Wu, Bing-Feng Lin, Hsien-Chen Chiu, Charlie Chung-Ping Chen |
ISCAS | 7 |
| 2011 | A 12 Gb/s chip-to-chip AC coupled transceiverabstractA differential AC coupled transceiver for high-speed and low-swing has been implemented in a 0.18μm CMOS process. The proposed architecture includes a pulse receiver and a broadband limiting amplifier to recover a NRZ signal from a 75fF capacitive coupled channel. The system works at 12Gb/s through 10cm FR4 printed circuit board interconnect, while dissipating only 13.5mW with a bit error rate less than 10-12. Yu-Shun Wang, Min-Han Hsieh, Yi-Chi Wu, Chia-Ming Liu, Hsien-Chen Chiu, Bing-Feng Lin, Charlie Chung-Ping Chen |
ISCAS | 7 |
| 2010 | Runtime temperature-based power estimation for optimizing throughput of thermal-constrained multi-core processorsabstractTechnology scaling has allowed integration of multiple cores into a single die. However, high power consumption of each core leads to very high heat density, limiting the throughput of thermal-constrained multi-core processors. To maximize the throughput, various software-based dynamic thermal management and optimization techniques have been proposed, many of which depend on accurate temperature sensing of each core. However, the decision for dynamic thermal management and throughput optimization only based on the temperature of each core can result in less optimal throughput in certain circumstances according to our investigation. In this paper, we propose 1) a dynamic power estimation method using a single thermal sensor for each core in multi-core processors, 2) a die temperature reconstruction method using the estimated power, and 3) a throughput optimization method based the estimated power instead of the temperature. According to our experiment using 90nm technology, the proposed method results in less than 3% error in estimating power and hot-spot temperature of a multi-core processor. Furthermore, the proposed throughput optimization method based on the estimated power leads to up to 4% higher throughput than a temperature-based optimization method. Dongkeun Oh, Nam Sung Kim, Charlie Chung-Ping Chen, Azadeh Davoodi, Yu Hen Hu |
ASP-DAC | 3 |
| 2010 | Interconnect delay and slew metrics using the beta distributionabstractIntegrated circuit process technology is entering the ultra deep submicron era. At this level, interconnect structure becomes very stiff and the metal resistance shielding effects problem is more serious. Although several delay metrics have been proposed, they are inefficient and difficult to implement. Hence, we propose a new delay and slew metric for interconnect based on Beta distribution and which does not require a look-up table to be built. Our metrics are efficient and easy to implement; the overall standard deviation and error mean are smaller than in previous works. Jun-Kuei Zeng, Charlie Chung-Ping Chen |
DATE | 2 |
| 2010 | Accurate and Analytical Statistical Spatial Correlation Modeling Based on Singular Value Decomposition for VLSI DFM ApplicationsabstractWith the significant advancement of statistical timing and yield analysis algorithms, there is a strong need for accurate and analytical spatial correlation models. In this paper, we propose a novel spatial correlation modeling method that can not only capture the general spatial correlation relationship but also can generate highly accurate and analytical models. Our method, based on singular value decomposition, can generate sequences of polynomial weighted by the singular values. Experimental results from foundry measurement data show that our proposed approach is 3x accuracy improvement over several distance based spatial correlation modeling methods. Jui-Hsiang Liu, Ming-Feng Tsai, Lumdo Chen, Charlie Chung-Ping Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2008 | LTCC spiral inductor modeling, synthesis, and optimizationabstractIn RF/microwave circuit design, inductor design is one of the most difficult and time-consuming tasks due to the tedious trial-and-error optimization process to achieve the target specifications. This paper brings forward a fast and accurate spiral inductor synthesis method, which automatically generates physical layout of inductors according to electrical specifications. This method is based on the fusion of substrate-aware PEEC model with optimal nonlinear optimization engine. Calculated inductance and Q values show good agreement with industrial field solver and measurement results. Tuck Boon Chan, Hsinchia Lu, Jun-Kuei Zeng, Charlie Chung-Ping Chen |
ASP-DAC | 4 |
| 2008 | An optimal algorithm for sizing sequential circuits for industrial library based designsabstractIn this paper, we propose an optimal gate sizing and clock skew optimization algorithm for globally sizing synchronous sequential circuits. The number of constraints and variables in our formulation is linear with respect to the number of circuit components and hence our algorithm can efficiently find the optimal solution for industrial scale designs. To the best of our knowledge our method is the first exact gate sizing algorithm that can handle cyclic sequential circuits. Experimental results on industrial cell libraries demonstrate that our algorithm can yield an average of 12.6% improvement in the optimal clock period by combining clock skew optimization with gate sizing. For identical clock period, our algorithm can achieve an average of 11.3% area savings over a popular commercial synthesis tool. Sanghamitra Roy, Yu Hen Hu, Charlie Chung-Ping Chen, Shih-Pin Hung, Tse-Yu Chiang, Jiuan-Guei Tseng |
ASP-DAC | 3 |
| 2008 | Accurate and analytical statistical spatial correlation modeling for VLSI DFM applicationsabstractWith the significant advancement of statistical timing and yield analysis algorithms, there is a strong need for accurate and analytical spatial correlation models. In this paper, we propose a novel spatial correlation modeling method not only can capture the general spatial correlation relationship but also can generate highly accurate and analytical models. Our method, based on Singular Value Decomposition (SVD), can generate sequences of polynomial weighted by the singular values. Experimental results from foundry measurement data show that our proposed approach is 5X accuracy improvement over several distance based spatial correlation modeling methods. Jui-Hsiang Liu, Ming-Feng Tsai, Lumdo Chen, Charlie Chung-Ping Chen |
DAC | 4 |
| 2008 | Deep Submicron Interconnect Timing Model with Quadratic Random Variable AnalysisabstractShrinking feature sizes and process variations are of increasing concern in modern technology. It is urgent that we develop statistical interconnect timing models which are harmonious with the current trend in statistical timing analysis flow. Although statistical model order reduction techniques have been explored, the statistical interconnect timing model has not yet been fully analyzed. In this work, we develop a novel algorithm and its corresponding analysis for the statistical interconnect timing model, using second-order statistical variations to model the non-Gaussian distribution effects. As this model is fully congruous with current statistical static timing analysis with the canonical model and does not require any Monte Carlo simulation analysis, performance is greatly improved. Experimental results show that the proposed closed-form quadratic interconnect timing model is within 0.0046% error of the corresponding Monte Carlo simulation. Jun-Kuei Zeng, Charlie Chung-Ping Chen |
DATE | 2 |
| 2008 | Performance measurement and queueing analysis of medium-high blocking probability of two and three parallel connection serversabstractIn this paper we propose a performance measurement and queueing analysis for medium-high blocking probability where the service rate is less than equal to the arrival rate of parallel connection servers, and for which we can estimate the system response time. First, we calculate the system response time of the parallel connection servers. Second, we simulate the queueing model of the parallel connection servers and derive the system response time. Third, we measure the system response time by using different numbers of ASP multiplication loops to represent the different service rates of the parallel connection servers. Fourth, we compare the simulation and measurement results with the medium-high blocking probability of the parallel connection from the different service rates of the servers. Fifth, we compare the system response time of both multi-parallel connection servers and a single server. Charlie Chung-Ping Chen, Ying-Wen Bai, Yin-Sheng Lee |
LCN | 1 |
| 2007 | SmartSmooth: A linear time convexity preserving smoothing algorithm for numerically convex data with application to VLSI designabstractConvex optimization problems are very popular in the VLSI design society due to their guaranteed convergence to a global optimal point. While optimizing tabular data, significant fitting efforts are required to fit the data into convex form. Fitting the tables into analytically convex forms like posynomials, suffers from excessive fitting errors, as the fitting problem may be non-convex. In recent literature optimal numerically convex tables have been proposed. Since these tables are numerical, it is extremely important to make the table data smooth, and yet preserve its convexity. The smoothness ensures that the convex optimizer behaves predictably and converges quickly to the global optimal point. The existing smoothing techniques either cannot preserve convexity, or require very high execution time. In this paper, we propose a linear time algorithm to smoothen a given numerically convex data and at the same time preserve convexity. Our proposed algorithm SmartSmooth can smoothen the data in linear time without introducing any additional error on the numerically convex data. We present our SmartSmooth results on industrial cell libraries. SmartSmooth when applied on convex tables produced by ConvexFit shows a 30times reduction in fitting square error over a posynomial fitting algorithm. Sanghamitra Roy, Charlie Chung-Ping Chen |
ASP-DAC | 2 |
| 2007 | Numerically Convex Forms and Their Application in Gate SizingabstractConvex-optimization techniques are very popular in the very large-scale-integration design society due to their guaranteed convergence to a global optimal point. The table data need to be fitted into convex forms to be used in the convex optimization problems. Fitting the tables into polynomials, which are analytically convex under logarithmic transformation, may suffer from the excessive fitting errors as the fitting problem is nonconvex. In this paper, we propose to directly adjust the lookup-table values into a numerically convex lookup table without any explicit analytical form. We show that numerically "convexifying" the lookup-table data with minimum perturbation can be formulated as a convex semidefinite optimization problem, and hence, optimality can be reached in polynomial time. We also propose three algorithms to make the table data smooth to enable faster convergence of the convex optimizer. Results from extensive experiments on industrial cell libraries demonstrate 9.6 improvement in fitting error over a well-developed polynomial-fitting procedure. We illustrate the effectiveness of this model in a convex optimization problem by providing results for using our model in the optimal gate sizing of standard cells. We observe a 5.07% improvement in the delay of International Symposium on Circuits and Systems (ISCAS) benchmark circuits over the polynomial-fitting procedure. Sanghamitra Roy, Weijen Chen, Charlie Chung-Ping Chen, Yu Hen Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2006 | Convergence-provable statistical timing analysis with level-sensitive latches and feedback loopsabstractStatistical timing analysis has been widely applied to predict the timing yield of VLSI circuits when process variations become significant. Existing statistical latch timing methods are either having exponential complexity or unable to treat the random variable's self-dependence caused by the coexistence of level-sensitive latches and feedback loops. In this paper, an efficient iterative statistical timing algorithm with provable convergence is proposed for latch-based circuits with feedback loops. Based on a new notion of iteration mean, we prove that the algorithm converges unconditionally. Moreover, we show that the converged value of iteration mean ca be used to predict the circuit yield during design time. Tested by ISCAS'89 benchmark circuits, the proposed algorithm shows a error of 1.1% and speedup of 303 /spl times/ on average when compared with the Monte Carlo simulation. Lizheng Zhang, Jeng-Liang Tsai, Weijen Chen, Yu Hen Hu, Charlie Chung-Ping Chen |
ASP-DAC | 5 |
| 2006 | Statistical timing analysis with path reconvergence and spatial correlationsabstractState of the art statistical timing analysis (STA) tools often yield less accurate results when timing variables become correlated. Spatial correlation and correlation caused by path reconvergence are among those which are most difficult to deal with. Existing methods treating these correlations will either suffer from high computational complexity or significant errors. In this paper, we present a sensitivity pruning method which significantly reduces the computational cost to consider path reconvergence correlation. We also develop an accurate and efficient model to deal with the spatial correlation. Lizheng Zhang, Yu Hen Hu, Charlie Chung-Ping Chen |
DATE | 3 |
| 2006 | Simultaneous area minimization and decaps insertion for power delivery network using adjoint sensitivity analysis with IEKS methodabstractThe soaring clocking frequency and integration density demand robust and stable power delivery to support tens of millions of transistor switching. In this paper, we consider the problem of minimizing the area of wires and decoupling capacitors (decaps) for a power delivery network, subject to the limit on integral of voltage drops. First, we derive the gradients of constraint function without Tellegen's theorem. This greatly simplifies the discuss of adjoint sensitivity analysis. Then, we apply the IEKS method to speed up the sensitivity analysis over 3 times. Finally, this efficient analyzer is incorporated with the state-of-the-art nonlinear programming package, SNOPT, to perform the optimization. Extensive experimental results show that the proposed method can work efficiently for large power delivery networks. Pei-Yu Huang, Yu-Min Lee, Jeng-Liang Tsai, Charlie Chung-Ping Chen |
ISCAS | 4 |
| 2006 | Non-gaussian statistical parameter modeling for SSTA with confidence interval analysisabstractMost of the existing statistical static timing analysis (SSTA) algorithms assume that the process parameters of have been given with 100% confidence level or zero errors and are preferable Gaussian distributions. These assumptions are actually quite questionable and require careful attention.In this paper, we aim at providing solid statistical analysis methods to analyze the measurement data on testing chips and extract the statistical distribution, either Gaussian or non-Gaussian which could be used in advanced SSTA algorithms for confidence interval or error bound information.Two contributions are achieved by this paper. First, we develop a moment matching based quadratic function modeling method to fit the first three moments of given measurement data in plain form which may not follow Gaussian distributions. Second, we provide a systematic way to analyze the confident intervals on our modeling strategies. The confidence intervals analysis gives the solid guidelines for testing chip data collections. Extensive experimental results demonstrate the accuracy of our algorithm. Lizheng Zhang, Charlie Chung-Ping Chen |
ISPD | 3 |
| 2006 | Temperature-Aware Placement for SOCsabstractDramatic rises in the power consumption and integration density of contemporary systems-on-chip (SoCs) have led to the need for careful attention to chip-level thermal integrity. High temperatures or uneven temperature distributions may result not only in reliability issues, but also timing failures, due to the temperature-dependent nature of chip time-to-failure and delay, respectively. To resolve these issues, high-quality, accurate thermal modeling and analysis, and thermally oriented placement optimizations, are essential prior to tapeout. This paper first presents an overview of thermal modeling and simulation methods, such as finite-difference time domain, finite element, model reduction, random walk, and Green-function based algorithms, that are appropriate for use in placement algorithms. Next, two-dimensional and three-dimensional thermal-aware placement algorithms such as matrix-synthesis, simulated annealing, partition-driven, and force directed are presented. Finally, future trends and challenges are described Jeng-Liang Tsai, Charlie Chung-Ping Chen, Brent Goplen, Haifeng Qian, Yong Zhan, Martin D. F. Wong, Sachin S. Sapatnekar |
Proc. IEEE | 2 |
| 2006 | Statistical static timing analysis with conditional linear MAX/MIN approximation and extended canonical timing modelabstractAn efficient and accurate statistical static timing analysis (SSTA) algorithm is reported in this paper, which features 1) a conditional linear approximation method of the MAX/MIN timing operator, 2) an extended canonical representation of correlated timing variables, and 3) a variation pruning method that facilitates intelligent tradeoff between simulation time and accuracy of simulation result. A special design focus of the proposed algorithm is on the propagation of the statistical correlation among timing variables through nonlinear circuit elements. The proposed algorithm distinguishes itself from existing block-based SSTA algorithms in that it not only deals with correlations due to dependence on global variation factors but also correlations due to signal propagation path reconvergence. Tested with the International Symposium on Circuits and Systems (ISCAS) benchmark suites, the proposed algorithm has demonstrated very satisfactory performance in terms of both accuracy and running time. Compared with Monte-Carlo-based statistical timing simulation, the output probability distribution got from the proposed algorithm is within 1.5% estimation error while a 350 times speed-up is achieved over a circuit with 5355 gates. Lizheng Zhang, Weijen Chen, Yu Hen Hu, Charlie Chung-Ping Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2006 | Correlation-Preserved Statistical Timing With a Quadratic Form of Gaussian VariablesabstractA recent study shows that the existing first-order canonical timing model is not sufficient to represent the dependency of the gate/wire delay on the processing and operational variations when these variations become more and more significant. Due to nonlinear mapping from variation sources to the gate/wire delay, the distribution of the delay will no longer be Gaussian even if variation sources are normally distributed. A novel “quadratic timing model” is proposed to capture the nonlinearity of the dependency of gate/wire delays and arrival times on the variation sources. Systematic methodology is also developed to evaluate the correlation and distribution of the quadratic timing model. Based on these, a statistical static timing analysis algorithm that retains the complete correlation information during timing analysis and has linear computation complexity with respect to both the circuit size and the number of variation sources is proposed. Tested on the ISCAS circuits, the proposed algorithm shows significant accuracy improvement over the existing first-order algorithm with a small amount of computational cost. Lizheng Zhang, Weijen Chen, Yu Hen Hu, John A. Gubner, Charlie Chung-Ping Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2005 | Fast and effective gate-sizing with multiple-Vt assignment using generalized Lagrangian RelaxationabstractSimultaneous gate-sizing with multiple V/sub t/ assignment for delay and power optimization is a complicated task in modern custom designs. In this work, we make the key contribution of a novel gate-sizing and multi-V/sub t/ assignment technique based on generalized Lagrangian relaxation. Experimental results show that our technique exhibits linear runtime and memory usage, and can effectively tune circuits with over 15,000 variables and 8,000 constraints in under 8 minutes (250/spl times/ faster than state-of-the-art optimization solvers). Hsinwei Chou, Yu-Hao Wang, Charlie Chung-Ping Chen |
ASP-DAC | 3 |
| 2005 | Comprehensive frequency dependent interconnect extraction and evaluation methodologyabstractThis paper presents a wide frequency range interconnect extraction and analysis methodology. First, an improved reluctance-based extraction algorithm is proposed to generate compact interconnect models at some sample frequencies. Then, DLSCF (Discrete Least Square Curve Fitting) techniques are employed to produce approximation polynomials to calculate parasitics at other frequencies. Finally, after transferring those approximation polynomials into power series of s and substituting them into the MNA (Modified Nodal Analysis) formula, we develop and apply the WIFRIM (Wide Frequency Range Interconnect Moment Matching) algorithm to calculate moments of arbitrary orders. Since WIFRIM only needs to decompose a sparse conductance matrix once, it results in significant speedup while providing accuracy within 1% error. Rong Jiang 0002, Charlie Chung-Ping Chen |
ASP-DAC | 2 |
| 2005 | Process-variation robust and low-power zero-skew buffered clock-tree synthesis using projected scan-line samplingabstractZero-skew clock-tree with minimum clock-delay is preferable due to its low unintentional and process-variation induced skews. We propose a zero-skew buffered clock-tree synthesis flow and a novel algorithm that enables clock-tree optimization throughout the full zero-skew design-space by considering simultaneous buffer-insertion, buffer-sizing, and wire-sizing. For an industrial clock-tree with 3101 sink nodes, our algorithm achieves up to 45X clock-delay improvement and up to 23% power reduction compared with its initial routing. Jeng-Liang Tsai, Charlie Chung-Ping Chen |
ASP-DAC | 2 |
| 2005 | Wave-pipelined on-chip global interconnectabstractA novel wave-pipelined global interconnect system is developed for reliable, high throughput, on-chip data communication. We argue that because there is only a single signal propagation path and a single type of 1-input gate(inverter), a wave-pipelined interconnect will have less stringent timing constraints than a wave-pipelined combinational logic block. A phase-lock loop based clock and data recovery unit architecture, adopted from off-chip high speed digital serial link, is designed for on-chip application so as to minimize power and area cost. Preliminary Monte Carlo simulation indicated that the wave-pipelined global interconnect architecture potentially can offer 18% higher throughput than a flip-flop pipelined global interconnect architecture at about the same level of reliability. While delivering data through long interconnect at the same bit rate, the wave-pipelined architecture consumes less power and requires less chip real estate. Lizheng Zhang, Yu Hen Hu, Charlie Chung-Ping Chen |
ASP-DAC | 3 |
| 2005 | Block based statistical timing analysis with extended canonical timing modelabstractBlock based statistical timing analysis (STA) tools often yield less accurate results when timing variables become correlated due to global source of variations and path reconvergence. To the best of our knowledge, no good solution is available handling both types of correlations simultaneously.In this paper, we present a novel statistical timing algorithm, AMECT (Asymptotic MAX/MIN approximation & Extended Canonical Timing model), that produces accurate timing estimation by handling both types of correlations simultaneously. An extended canonical timing model is developed to evaluate and decompose correlations between arbitrary timing variables. And an intelligent pruning method is designed enabling trade-off runtime with accuracy.Tested with ISCAS benchmark suites, AMECT shows both high accuracy and high performance compared with Monte Carlo simulation results: with distribution estimation error < 1.5% while with around 350X speed up on a circuit with 5355 gates. Lizheng Zhang, Yu Hen Hu, Charlie Chung-Ping Chen |
ASP-DAC | 3 |
| 2005 | ICCAP: a linear time sparse transformation and reordering algorithm for 3D BEM capacitance extractionabstractThis paper presents an efficient hierarchical 3D capacitance extraction algorithm - ICCAP. Most previous capacitance extraction algorithms introduce intermediate variables to facilitate the hierarchical potential calculation but still preserve the leaf panels as the basis. In this paper, we discover that those intermediate variables are fundamentally much better basis than leaf panels. As a result, we are able to explicitly construct the sparse potential coefficient matrix and solve it with linear memory in linear runtime. Furthermore, the explicit sparse formulation not only enables the usage of preconditioned iterative Krylov subspace methods but also the reordering technique. A new reordering technique is proposed to further reduce over 20% of memory consumption and runtime in comparison to no reordering techniques applied. Experimental results demonstrate the superior runtime and memory consumption of ICCAP over previous approaches while achieving similar accuracy. Rong Jiang 0002, Yi-Hao Chang, Charlie Chung-Ping Chen |
DAC | 3 |
| 2005 | Correlation-preserved non-gaussian statistical timing analysis with quadratic timing modelabstractRecent study shows that the existing first order canonical timing model is not sufficient to represent the dependency of the gate delay on the variation sources when processing and operational variations become more and more significant. Due to the nonlinearity of the mapping from variation sources to the gate/wire delay, the distribution of the delay is no longer Gaussian even if the variation sources are normally distributed. Anovelquadratic timing model is proposed to capture the non-linearity of the dependency of gate/wire delays and arrival times on the variation sources. Systematic methodology is also developed to evaluate the correlation and distribution of the quadratic timing model. Based on these, a novel statistical timing analysis algorithm is propose which retains the complete correlation information during timing analysis and has the same computation complexity as the algorithm based on the canonical timing model. Tested on the ISCAS circuits, the proposed algorithm shows 10 × accuracy improvement over the existing first order algorithm while no significant extra runtime is needed. Lizheng Zhang, Weijen Chen, Yu Hen Hu, John A. Gubner, Charlie Chung-Ping Chen |
DAC | 5 |
| 2005 | Statistical Timing Analysis with Extended Pseudo-Canonical Timing ModelabstractState of the art statistical timing analysis (STA) tools often yield less accurate results when timing variables become correlated due to global source of variations and path reconvergence. To the best of our knowledge, no good solution is available for dealing both types of correlations simultaneously. In this paper, we present a novel extended pseudo-canonical timing model to retain and evaluate both types of correlation during statistical timing analysis with minimum computation cost. Also, an intelligent pruning method is introduced to enable trade-off runtime with accuracy. Tested with ISCAS benchmark suites, our method shows both high accuracy and high performance. For example, on the circuit c6288, our distribution estimation error shows 15/spl times/ accuracy improvement compared with previous approaches. Lizheng Zhang, Weijen Chen, Yu Hen Hu, Charlie Chung-Ping Chen |
DATE | 4 |
| 2005 | 1-V 7-mW dual-band fast-locked frequency synthesizerabstractThis paper presents a fully integrated 1-V, dual band, fast-locked frequency synthesizer for IEEE 802.11 a/b/g WLAN applications. It can synthesize frequencies in the range of 2.4 - 2.7 GHz with a step of 9.375 MHz, and in the range of 5.14 - 5.70 GHz with a step of 20 MHz. Simulation using 0.18-μm rf and mixed-signal CMOS technology demonstrates a total power consumption of 7-mW. An adaptive bandwidth controller is employed to achieve a fast locking time. The frequency divider combines the conventional and the extended true-single-phase-clock logics. To ensure a proper dividing function, a cascode voltage switch (CVS) topology is used in the preamplifier stage. The reference spurs at an offset of 10-MHz are as low as -80 dBc, and the phase noise at an offset of 1 MHz is lower than -118 dBc for the entire tuning range. Chien-Liang Chen, Charlie Chung-Ping Chen |
ACM Great Lakes Symposium on VLSI | 3 |
| 2005 | Efficient statistical capacitance variability modeling with orthogonal principle factor analysisabstractDue to the ever-increasing complexity of VLSI designs and IC process technologies, the mismatch between a circuit fabricated on the wafer and the one designed in the layout tool grows ever larger. Therefore, characterizing and modeling process variations of interconnect geometry has become an integral part of analysis and optimization of modern VLSI designs. In this paper, we present a systematic methodology to develop a closed form capacitance model, which accurately captures the nonlinear relationship between parasitic capacitances and dominant global/local process variation parameters. The explicit capacitance representation applies the orthogonal principle factor analysis to greatly reduce the number of random variables associated with modeling conductor surface fluctuations while preserving the dominant sources of variations, and consequently the variational capacitance model can be efficiently utilized by statistical model order reduction and timing analysis tools. Experimental results demonstrate that the proposed method exhibits over 100/spl times/ speedup compared with Monte Carlo simulation while having the advantage of generating explicit variational parasitic capacitance models of high order accuracy. Rong Jiang 0002, Wenyin Fu, Janet Roveda, Vince Lin, Charlie Chung-Ping Chen |
ICCAD | 5 |
| 2005 | System-level power and thermal modeling and analysis by orthogonal polynomial based response surface approach (OPRS)abstractThis paper proposes a new statistical response surface based power estimation technique. The new approach is able to include a number of parameters such as multiple Vdd, multiple Vth and gate sizing parameters. It has both deterministic ability and statistical ability. The deterministic ability allows the new model to provide optimal design parameters for power reduction. The statistical ability can be used to model the process variation impact on power. Janet Roveda, Bharat Srinivas, Charlie Chung-Ping Chen, Jun Li 0066 |
ICCAD | 4 |
| 2005 | EPEEC: comprehensive SPICE-compatible reluctance extraction for high-speed interconnects above lossy multilayer substratesabstractWith continuous advances in radio-frequency (RF) mixed-signal very large scale integration (VLSI) technology, the creation of eddy currents in lossy multilayer substrates has made the already complicated interconnect analysis and modeling issue more challenging. To account for substrate losses, traditional electromagnetic methods are often computationally prohibitive for today's VLSI geometries. In this paper, an accurate and efficient interconnect modeling approach-the eddy-current-aware partial equivalent element circuit (EPEEC)-is proposed. Based on complex image theory, it extends the traditional partial equivalent element circuit (PEEC) model to simultaneously take multilayer substrate eddy-current losses and frequency-dependent effects into consideration. To accommodate even larger scale on-chip interconnect networks, EPEEC develops a new simulation program with integrated circuit emphasis (SPICE)-compatible reluctance extraction algorithm by applying sparsification in the inverse inductance domain with an extended window algorithm. Compared with several industry standard inductance and full-wave solvers, such as FastHenry and Sonnet, EPEEC demonstrates within 1.5% accuracy while providing over 100/spl times/ speedup. Rong Jiang 0002, Wenyin Fu, Charlie Chung-Ping Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2005 | HiPRIME: hierarchical and passivity preserved interconnect macromodeling engine for RLKC power deliveryabstractThis paper proposes a general hierarchical analysis methodology, HiPRIME, to efficiently analyze RLKC power delivery systems. After partitioning the circuits into blocks, we develop and apply the IEKS (Improved Extended Krylov Subspace) method to build the multiport Norton equivalent circuits which transform all the internal sources to Norton current sources at ports. Since there are no active elements inside the Norton circuits, passive or realizable model order reduction techniques such as PRIMA can be applied. The significant speed improvement, 700 times faster than Spice with less than 0.2% error and 7 times faster than a state-of-the-art solver, InductWise, is observed. To further reduce the top-level hierarchy runtime, we develop a second-level model reduction algorithm and prove its passivity. Yu-Min Lee, Yahong Cao, Tsung-Hao Chen, Janet Roveda, Charlie Chung-Ping Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2004 | Frequency-dependent reluctance extraction
Clement Luk, Tsung-Hao Chen, Charlie Chung-Ping Chen |
ASP-DAC | 3 |
| 2004 | Statistical timing analysis in sequential circuit for on-chip global interconnect pipeliningabstractWith deep-sub-micron (DSM) technology, statistical timing analysis becomes increasingly crucial to characterize signal transmission over global interconnect wires. In this paper, a novel statistical timing analysis approach has been developed to analyze the behavior of two important pipelined architectures for multiple clock-cycle global interconnect, namely, the flip-flop inserted global wire and the latch inserted global wire. We present analytical formula that is based on parameters obtained using Monte Carlo simulation. These results enable a global interconnect designer to explore design trade-offs between clock frequency and probability of bit-error during data transmission. Categories and Subject Descriptors Lizheng Zhang, Yu Hen Hu, Charlie Chung-Ping Chen |
DAC | 3 |
| 2004 | SCORE: SPICE COmpatible Reluctance ExtractionabstractPresently, a necessary modification to mainstream analysis tools prevents the direct application of reluctance k. In this paper, we propose a reluctance realization algorithm (RRA) by directly converting reluctances to circuit elements compatible with general simulation engines, such as SPICE. Reluctance realization is applicable to arbitrary circuit topology and no accuracy penalty is involved in the realization process. Since the stability of the converted circuit largely depends on the stability of the reluctance matrix, we present an efficient improved recursive bisection cutting algorithm (IRBCA) to obtain stability-guaranteed reluctance matrices, and integrate IRBCA and RRA into a SPICE compatible reluctance extraction tool, SCORE. Rong Jiang 0002, Charlie Chung-Ping Chen |
DATE | 2 |
| 2004 | Realizable Reduction for Electromagnetically Coupled RLMC InterconnectsabstractThis paper presents a realizable RLMC reduction algorithm for extracted interconnect circuits based on two effective approaches: RL branch reduction and RC/LC node reduction. Our algorithm takes advantage of some structures existing extensively in interconnect circuits and hence has extremely fast execution time. It takes about 8 seconds to reduce a circuit of over 300,000 elements while maintaining 3% error and 75% element reduction ratio. Rong Jiang 0002, Charlie Chung-Ping Chen |
DATE | 2 |
| 2004 | Thermal and Power Integrity Based Power/Ground Networks OptimizationabstractWith the increasing power density and heat-dissipation cost of modern VLSI designs, thermal and power integrity has become serious concern. Although the impacts of thermal effects on transistor and interconnect performance are well-studied, the interactions between power-delivery and thermal effects are not clear. As a result, power-delivery design without thermal consideration may cause soft-error, reliability degradation, and even premature chip failures. In this paper, we propose a thermal-aware power-delivery optimization algorithm. By simultaneously considering thermal and power integrity, we are able to achieve high power supply quality and thermal reliability. For a 58/spl times/72 mesh as shown in the experimental results, our algorithm shows that the lifetime of the optimized ground network is 9.5 years. Whereas the lifetime of the ground network generated by a traditional method is only 2 years without thermal concern. Ting-Yuan Wang, Jeng-Liang Tsai, Charlie Chung-Ping Chen |
DATE | 3 |
| 2004 | A yield improvement methodology using pre- and post-silicon statistical clock schedulingabstractIn deep sub-micron technologies, process variations can cause significant path delay and clock skew uncertainties thereby lead to timing failure and yield loss. In this paper, we propose a comprehensive clock scheduling methodology that improves timing and yield through both pre-silicon clock scheduling and post-silicon clock tuning. First, an optimal clock scheduling algorithm has been developed to allocate the slack for each path according to its timing uncertainty. To balance the skew that can be caused by process variations, programmable delay elements are inserted at the clock inputs of a small set of flip-flops on the timing critical paths. A delay-fault testing scheme combined with linear programming is used to identify and eliminate timing violations in the manufactured chips. Experimental results show that our methodology achieves substantial yield improvement over a traditional clock scheduling algorithm in many of the ISCAS89 benchmark circuits, and obtain an average yield improvement of 13.6%. Jeng-Liang Tsai, Dong Hyun Baik, Charlie Chung-Ping Chen, Kewal K. Saluja |
ICCAD | 3 |
| 2004 | Sensitivity guided net weighting for placement driven synthesisabstractNet weighting is a key technique in large scale timing driven placement, which plays a crucial role for deep submicron physical synthesis and timing closure. A popular way to assign net weight is based on the slack of the nets, trying to minimize the worst negative slack (WNS) for the entire circuit. While WNS is an important optimization metric, another figure of merit (FOM), defined as the total slack difference compared to a certain slack threshold for all timing end points, is of equivalent importance to measure the overall timing closure result for highly complex modern ASIC and microprocessor designs. In this paper, we perform a comprehensive analysis of the slack and FOM sensitivities to the net weight, and propose a new net weighting scheme based on the slack and FOM sensitivities. Such sensitivity analysis implicitly takes potential physical synthesis effect into consideration. Experiment results on a set of industrial circuits are promising for both stand-alone timing driven placement and physical synthesis afterwards. Ting-Yuan Wang, Jeng-Liang Tsai, Charlie Chung-Ping Chen |
ISPD | 3 |
| 2004 | Zero skew clock-tree optimization with buffer insertion/sizing and wire sizingabstractClock distribution is crucial for timing and design convergence in high-performance very large scale integration designs. Minimum-delay/power zero skew buffer insertion/sizing and wire-sizing problems have long been considered intractable. In this paper, we present ClockTune , a simultaneous buffer insertion/sizing and wire-sizing algorithm which guarantees zero skew and minimizes delay and power in polynomial time. Extensive experimental results show that our algorithm executes very efficiently. For example, ClockTune achieves 45/spl times/ delay improvement for buffering and sizing an industrial clock tree with 3101 sink nodes on a 1.2-GHz Pentium IV PC in 16 min, compared with the initial routing. Our algorithm can also be used to achieve useful clock skew to facilitate timing convergence and to incrementally adjust the clock tree for design convergence and explore delay-power tradeoffs during design cycles. ClockTune is available on the web (http://vlsi.ece.wisc.edu/Tools.htm). Jeng-Liang Tsai, Tsung-Hao Chen, Charlie Chung-Ping Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2003 | A hierarchical analysis methodology for chip-level power delivery with realizable model reductionabstractAbstract — In this paper, we propose a novel hierarchical analysis methodology to facilitate efficient chip-level power fluctuation analysis. With extreme efficiency and simplicity, our design methodology first builds time-varying multiport Norton equivalent circuits in a row-by-row or block-by-block based followed by global analysis on the integrated reduced models. After generating the Norton equivalent sources at external ports, we apply realizable model order reduction technologies to further reduce model. Since the elements of our reduced model are also RC devices, they are fully compatible with general circuit simulation engines. The experimental results demonstrate more than 4X speed up with the flat simulation while maintaining within 5 % of accuracy. I. Yu-Min Lee, Charlie Chung-Ping Chen |
ASP-DAC | 2 |
| 2003 | The Power Grid Transient Simulation in Linear Time Based on 3D Alternating-Direction-Implicit Method
Yu-Min Lee, Charlie Chung-Ping Chen |
DATE | 2 |
| 2003 | SuPREME: Substrate and Power-delivery Reluctance-Enhanced Macromodel Evaluation
Tsung-Hao Chen, Clement Luk, Charlie Chung-Ping Chen |
ICCAD | 3 |
| 2003 | Optimal minimum-delay/area zero-skew clock tree wire-sizing in pseudo-polynomial timeabstractIn 21st-Century VLSI design, clocking plays crucial roles for both performance and timing convergence. Due to their non-convex nature, optimal minimum-delay/area zero-skew wire-sizing problems have long been considered intractable. None of the existing approaches can guarantee optimality for general clock trees to the authors' best knowledge. In this paper, we present an ε-optimal zero-skew wire-sizing algorithm, ClockTune, which guarantees zero-skew with delay and area within ε distance to the optimal solutions in pseudo-polynomial time. Extensive experimental results show that our algorithm executes very efficiently in both runtime and memory usage. For example, ClockTune takes less than two minutes and 35MB memory to size an industrial clock tree with 3101 sink nodes within 2% to the optimal solution on a 533MHz Pentium III PC. Our algorithm can also be used to achieve useful clock skew to facilitate timing convergence and to incrementally adjust clock tree for design convergence and explore delay/power tradeoffs during design cycles. ClockTune is available on the web [13]. Jeng-Liang Tsai, Tsung-Hao Chen, Charlie Chung-Ping Chen |
ISPD | 3 |
| 2003 | 3D thermal-ADI: an efficient chip-level transient thermal simulatorabstractRecent studies show that the nonuniform thermal distribution on the substrate and interconnects has impact on the circuit reliability and performance. Hence three-dimensional (3-D) thermal analysis is crucial to analyze these effects. In this paper, we present and develop an efficient 3-D transient thermal simulator based on the alternating direction implicit (ADI) method for large scale temperature estimation problems. Our simulator, 3D Thermal-ADI, not only has a linear runtime and memory requirement, but also is unconditionally stable. Detailed analysis of the 3-D nonhomogeneous cases and boundary conditions for on-chip VLSI applications are introduced and presented. Extensive experimental results show that our algorithm is not only orders of magnitude faster than the traditional thermal simulation algorithms, but is also highly accurate and memory efficient. The temperature profile of steady state can be reached in few iterations. The software is avaiable on the web [1]. Ting-Yuan Wang, Yu-Min Lee, Charlie Chung-Ping Chen |
ISPD | 3 |
| 2003 | INDUCTWISE: inductance-wise interconnect simulator and extractorabstractA robust, efficient, and accurate inductance extraction and simulation tool, INDUCTWISE, is developed and described in this paper. This work advances the state-of-the-art inductance extraction and simulation techniques, and has several major contributions. First, albeit the great benefits of efficiency, the recently proposed inductance matrix sparsification algorithm, the-method (Ji et al. 2001), has a flaw in the stability proof for general geometry. We provide a theoretical analysis as well as a provable stable algorithm for it. Second, a robust window-selection algorithm is presented for general geometry. Third, integrated with the nodal analysis formulation, INDUCTWISE achieves exceptional performance without frequency-dependent complex operations and directly gives time-domain responses. Experimental results show that INDUCTWISE extractor and simulator have dramatic speedup compared to FastHenry and SPICE3, respectively. It has been well tested and released on the web for public usage (Available: http://vlsi.ece.wisc.edu/Inductwise.htm). Tsung-Hao Chen, Clement Luk, Charlie Chung-Ping Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2003 | The power grid transient simulation in linear time based on 3-D alternating-direction-implicit methodabstractThe rising power consumption and clock frequency of very large scale integration technology demand robust and stable power delivery. Extensive transient simulations on large-scale power delivery structures are required to analyze power delivery fluctuation caused by dynamic IR drop and Ldi/dt drop as well as package and on-chip resonance. In this paper, we develop a novel and efficient transient simulation algorithm for the power distribution networks. Our algorithm, three-dimensional (3-D) transmission-line-modeling alternating-direction-implicit (TLM-ADI) method, first models the power delivery structure as 3-D transmission line shunt-node structure and transfers those equations to the telegraph equation. Finally, we solve it by the alternating direction implicit method. The 3-D TLM-ADI method, with linear runtime and memory requirement, is also unconditionally stable, which ensures that the time steps are not limited by any stability requirement. Extensive numerical simulation results show that the proposed algorithm is not only over 300 000 times faster than SPICE but also extremely memory saving and accurate. Yu-Min Lee, Charlie Chung-Ping Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2003 | Thermal-ADI - a linear-time chip-level dynamic thermal-simulation algorithm based on alternating-direction-implicit (ADI) methodabstractDue to the dramatic increase of clock frequency and integration density, power density and on-chip temperature in high-end very large scale integration (VLSI) circuits rise significantly. To ensure the timing correctness and the reliability of high-end VLSI design, efficient and accurate chip-level transient thermal simulations are of crucial importance. In this paper, we develop and present an efficient transient thermal-simulation algorithm based on the alternating-direction-implicit (ADI) method. Our algorithm, thermal-ADI, not only has a linear run time and memory requirement , but is also unconditionally stable, which ensures that time step is not limited by any stability requirement. Extensive experimental results show that our algorithm is not only orders of magnitude faster than the traditional thermal-simulation algorithms, but also highly accurate and efficient in memory usage. Ting-Yuan Wang, Charlie Chung-Ping Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2002 | HiPRIME: hierarchical and passivity reserved interconnect macromodeling engine for RLKC power deliveryabstractThis paper proposes a general hierarchical analysis methodology, HiPRIME, to efficiently analyze RLKC power delivery systems. After partitioning the circuits into blocks, we develop and apply the IEKS (Improved Extended Krylov Subspace) method to build the Multi-port Norton Equivalent circuits which transform all the internal sources to Norton current sources at ports. Since there is no active elements inside the Norton circuits, passive or realizable model order reduction techniques such as PRIMA can be applied. To further reduce the top-level hierarchy runtime, we develop a second-level model reduction algorithm and prove its passivity. Experimental results show 400-700X runtime improvement with less than 0.2% error. Yahong Cao, Yu-Min Lee, Tsung-Hao Chen, Charlie Chung-Ping Chen |
DAC | 4 |
| 2002 | INDUCTWISE: inductance-wise interconnect simulator and extractorabstractWe develop a robust, efficient, and accurate tool, which integrates inductance extraction and simulation, called INDUCTWISE. This paper advances the state-of-the-art inductance extraction and simulation techniques and contains two major parts. In the first part, INDUCTWISE extractor, we discover the recently proposed inductance matrix sparsification algorithm, the K-method[1], albeit its great benefits of efficiency, has a major flaw on the stability. We provide both a counter example and a remedy for it. A window section algorithm is also presented to preserve the accuracy of the sparsification method. The second part, INDUCTWISE simulator, demonstrates great efficiency of integrating the nodal analysis formulation with the improved K-method. Experimental results show that INDUCTWISE has over 250x speedup compared to SPICE3. The proposed sparsification algorithm accelerates the simulator another 175x and speeds up the extractor 23.4x within 0.1% of error. INDUCTWISE can extract and simulate an 118K-conductor RKC circuit within 18 minutes. It has been well tested and released on the web for public usage. (http://vlsi.ece.wisc.edu/Inductwise.htm) Tsung-Hao Chen, Clement Luk, Hyungsuk Kim, Charlie Chung-Ping Chen |
ICCAD | 4 |
| 2002 | Power grid transient simulation in linear time based on transmission-line-modeling alternating-direction-implicit methodabstractThe soaring clocking frequency and integration density demand robust and stable power delivery to support tens of millions of transistors switching. To ensure the design quality of power delivery, extensive transient power grid simulations need to be performed during the design process. However, the traditional circuit simulation engines are not scaled well for the complexity of power delivery. As a result, it often takes a long runtime and huge memory requirement to simulate a medium-sized power grid circuit. In this paper, the authors develop and present a new efficient transient simulation algorithm for power distribution. The proposed. algorithm, transmission-line-modeling alternating-direction-implicit (TLM-ADI), first models the power delivery structure as transmission line mesh structure, then solves the transient modified nodal analysis matrices by the alternating-direction-implicit method. The proposed algorithm, with linear runtime and memory requirement, is also unconditionally stable which ensures that the time-step is not limited by any stability requirement. Extensive experimental, results show that the proposed algorithm is not only orders of magnitude faster than SPICE but also extremely memory saving and accurate. Yu-Min Lee, Charlie Chung-Ping Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2002 | 3-D Thermal-ADI: a linear-time chip level transient thermal simulatorabstractRecent study shows that the nonuniform thermal distribution not only has an impact on the substrate but also interconnects. Hence, three-dimensional (3-D) thermal analysis is crucial to analyze these effects. In this paper, the authors present and develop an efficient 3-D transient thermal simulator based on the alternating direction implicit (ADI) method for temperature estimation in a 3-D environment. Their simulator, 3D Thermal-ADI, not only has a linear runtime and memory requirement, but also is unconditionally stable. Detailed analysis of the 3-D nonhomogeneous cases and boundary conditions for on-chip VLSI applications are introduced and presented. Extensive experimental results show that our algorithm is not only orders of magnitude faster than the traditional thermal simulation algorithms but also highly accurate and memory efficient. The temperature profile of steady state can also be reached in several iterations. This software will be released via the web for public usage. Ting-Yuan Wang, Charlie Chung-Ping Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2001 | Optimal spacing and capacitance padding for general clock structuresabstractClock-tuning has been classified as important but tough tasks due to the non-convex nature caused by the skew requirements. As a result, all existing mathematical programming approaches are often trapped at local minimum and have no guarantee of obtaining global optimal solution. In this paper, we present optimal clock tuning algorithms which effectively apply capacitance-padding to reduce clock skew, power, and delay for general clock topologies. Capacitance-padding can be achieved by wire-spacing, wire-splitting, wire-padding and transistor-padding. We show that under the El-more delay model, capacitance-padding can be formulated as a linear programming problem and solved with great efficiency. Capacitance-padding can also be used as a post processing step for any non-zero-skew clock tree or mesh structure to achieve timing closure. Experiment results on several practical industry examples show that our algorithms are extremely efficient. Problems with over 6000 variables can be optimally tuned within 1 minute on a PC with 500 -MHZ Intel Pentium III processor. Yu-Min Lee, Hing Yin Lai, Charlie Chung-Ping Chen |
ASP-DAC | 3 |
| 2001 | Efficient Large-Scale Power Grid Analysis Based on Preconditioned Krylov-Subspace Iterative MethodsabstractIn this paper, we propose preconditioned Krylov-subspace iterative methods to perform efficient DC and transient simulations for large-scale linear circuits with an emphasis on power delivery circuits. We also prove that a circuit with inductors can be simplified from MNA to NA format, and the matrix becomes an s.p.d matrix. This property makes it suitable for the conjugate gradient with incomplete Cholesky decomposition as the preconditioner, which is faster than other direct and iterative methods. Extensive experimental results on large-scale industrial power grid circuits show that our method is over 200 times faster for DC analysis and around 10 times faster for transient simulation compared to SPICE3. Furthermore, our algorithm reduces over 75% of memory usage than SPICE3 while the accuracy is not compromised. Tsung-Hao Chen, Charlie Chung-Ping Chen |
DAC | 2 |
| 2001 | Hierarchical model order reduction for signal-integrity interconnect synthesisabstractArticle Share on Hierarchical model order reduction for signal-integrity interconnect synthesis Authors: Yu-Min Lee Department of Electrical and Computer Engineering, University of Wisconsin at Madison, Madison, WI Department of Electrical and Computer Engineering, University of Wisconsin at Madison, Madison, WIView Profile , Charlie Chung-Ping Chen Department of Electrical and Computer Engineering, University of Wisconsin at Madison, Madison, WI Department of Electrical and Computer Engineering, University of Wisconsin at Madison, Madison, WIView Profile Authors Info & Claims GLSVLSI '01: Proceedings of the 11th Great Lakes symposium on VLSIMarch 2001Pages 109–114https://doi.org/10.1145/368122.368883Published:01 March 2001Publication History 3citation178DownloadsMetricsTotal Citations3Total Downloads178Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Publisher SiteGet Access Yu-Min Lee, Charlie Chung-Ping Chen |
ACM Great Lakes Symposium on VLSI | 2 |
| 2001 | Power Grid Transient Simulation in Linear Time Based on Transmission-Line-Modeling Alternating-Direction-Implicit MethodabstractThe soaring clocking frequency and integration density demand robust and stable power delivery to support tens of millions of transistors switching. To ensure the design quality of power delivery, extensive transient power grid simulations need to be performed during design process. However, the traditional circuit simulation engines are not scaled as well as the complexity of power delivery, as a result, it often takes a long runtime and huge memory requirement to simulate a medium size power grid circuit. We develop and present a new efficient transient simulation algorithm for power distribution. The proposed algorithm, TLM-ADI (transmission-line-modeling alternatingdirection-implicit), first models the power delivery structure as transmission line mesh structure, then solves the transient MNA matrices by the alternating-direction-implicit method. The proposed algorithm, with linear runtime and memory requirement, is also unconditionally stable which ensures that the time-step is not limited by any stability requirement. Extensive experimental results show that the proposed algorithm is not only orders of magnitude faster than SPICE but also extremely accurate. Yu-Min Lee, Charlie Chung-Ping Chen |
ICCAD | 2 |
| 2001 | Linear Time Hierarchical Capacitance Extraction without Multipole ExpansionabstractHierarchical capacitance extraction algorithms have been shown an efficient and accurate capacitance extraction algorithm. An improved algorithm is also proposed to remove its runtime dependency on the number of conductors by a combination of hierarchical and multipole expansion algorithm. In this paper, we show that with the introduction of hierarchical merging operation and super-node representation, we can achieve linear runtime and accuracy without involving multipole expansion. Experimental results show over 10/spl times/ runtime improvement and 20/spl times/ memory saving over the multipole approaches with comparable accuracy and better numerical stability. Saisanthosh Balakrishnan, Hyungsuk Kim, Yu-Min Lee, Charlie Chung-Ping Chen |
ICCD | 5 |
| 2001 | RC-in RC-out Model Order Reduction Accurate up to Second Order MomentsabstractIn this paper, we present a RC-in RC-out model order reduction method which takes RC circuits and accurate reduced models which can be realized using passive RC elements. The reduced models are accurate up to 2nd order moment and hence are more accurate than the first order moment matching based algorithm. The runtime and reduction ratios of our method are not dependent on the number of ports, which can be very large for tightly coupled interconnects. Extensive SPICE simulations show the average accuracy of our algorithm is 5/spl times/ better than first order moment based algorithms. It also significantly improves the reduction ratio by 50% comparing the first order moment based algorithm for the same accuracy. Pradeepsunder Ganesh, Charlie Chung-Ping Chen |
ICCD | 2 |
| 2001 | Thermal-ADI: a linear-time chip-level dynamic thermal simulation algorithm based on alternating-direction-implicit (ADI) methodabstractDue to the dramatic increase of clock frequency and integration density, power density and on-chip temperature in high-end VLSI circuits rise significantly. To ensure the timing correctness and the reliability of high-end VLSI design, efficient and accurate chip-level transient thermal simulations are of crucial importance. Ting-Yuan Wang, Charlie Chung-Ping Chen |
ISPD | 2 |
| 2000 | Generalized FDTD-ADI: An Unconditionally Stable Full-Wave Maxwell's Equations Solver for VLSI Interconnect ModelingabstractThe finite-difference time-domain (FDTD) method of solving the full-wave Maxwell's equations has been recently extended to provide accurate and numerically stable operation for time steps exceeding the Courant limit. The elimination of an upper bound on the size of the time step was achieved using an alternating-implicit direction (ADI) time-stepping scheme. This greatly increases the computational efficiency of the FDTD method for classes of problems where the cell size of the three-dimensional space lattice is constrained to be much smaller than the shortest wavelength in the source spectrum. One such class of problems is the analysis of highspeed VLSI interconnects where full-wave methods are often needed for the accurate analysis of parasitic electromagnetic wave phenomena. In this paper, we present an enhanced FDTD-ADI formulation which permits the modeling of realistic lossy materials such as semiconductor substrates and metal conductors as well as artificial lossy materials needed for perfectly matched layer (PML) absorbing boundary conditions. Simulations using our generalized FDTD-ADI formulation are presented to demonstrate the accuracy and extent to which the computational burden is reduced by the ADI scheme. Charlie Chung-Ping Chen, Tae-Woo Lee, Narayanan Murugesan, Susan C. Hagness |
ICCAD | 1 |
| 1999 | Noise-Aware Repeater Insertion and Wire-Sizing for On-Chip Interconnect Using Hierarchical Moment-MatchingabstractRecently, several algorithms for interconnect optimization via repeater insertion and wire sizing have appeared based on the Elmore delay model. Using the Devgan noise metric [6] a noiseaware repeater insertion technique has also been proposed recently. Recognizing the conservatism of these delay and noise models, we propose a moment-matching based technique to interconnect optimization that allows for much higher accuracy while preserving the hierarchical nature of Elmore-delay-based techniques. We also present a novel approach to noise computation that accurately captures the effect of several attackers in linear time with respect to the number of attackers and wire segments. Our practical experiments with industrial nets indicate that the corresponding reduction in error afforded by these more accurate models justifies this increase in runtime for aggressive designs which is our targeted domain. Our algorithm yields delay and noise estimates within 5% of circuit simulation results. Charlie Chung-Ping Chen, Noel Menezes |
DAC | 1 |
| 1999 | Error Bounded Padé Approximation via Bilinear Conformal TransformationabstractSince Asymptotic Waveform Evaluation (AWE) was introduced in [5], many interconnect model order reduction methods via Pad& approximation have been proposed.Although the stability and precision of model reduction methods have been greatly improved, the following important question has not been answered: "What is the error bound in the time domain?".This problem is mainly caused by the "gap" between the frequency domain and the time domain, i.e. a good approximated transfer function in the frequency domain may not be a good approximation in the time domain.All of the existing methods approximate the transfer function directly in the frequency domain and hence can not provide error bounds in the time domain.In this paper, we present new moment matching methods which can provide guaranteed error bounds in the time domain.Our methods are based on the classic work by Teasdale in [l] which performs Pade approximation in a transformed domain by the bilinear conformal transformation s = E. Charlie Chung-Ping Chen, Martin D. F. Wong |
DAC | 1 |
| 1999 | Fast and exact simultaneous gate and wire sizing by Lagrangian relaxationabstractThis paper considers simultaneous gate and wire sizing for general very large scale integrated (VLSI) circuits under the Elmore delay model. We present a fast and exact algorithm which can minimize total area subject to maximum delay bound. The algorithm can be easily modified to give exact algorithms for optimizing several other objectives (e.g., minimizing maximum delay or minimizing total area subject to arrival time specifications at all inputs and outputs). No previous algorithm for simultaneous gate and wire sizing can guarantee exact solutions for general circuits. Our algorithm is an iterative one with a guarantee on convergence to global optimal solutions. It is based on Lagrangian relaxation and "one-gate/wire-at-a-time" greedy optimizations, and is extremely economical and fast. For example, we can optimize a circuit with 27648 gates and wires in 11.53 min using under 23 Mbytes memory on a PC with a 333-MHz Pentium II processor. Charlie Chung-Ping Chen, Chris C. N. Chu, Martin D. F. Wong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1998 | Fast and exact simultaneous gate and wire sizing by Lagrangian relaxationabstractThis paper considers simultaneous gate and wire sizing for general VUI circuits under the E[more delay model.JVepresent a fast and met algorithm which can minimize total area subject to mmimurn delay bound.The algorithm can be easily modl~ed to give ~wct algorithms for optimizing several other objectives (e.g.minimizing maximum delay or minimizing total area subject to arrival time specl~cations at all inputs and outputs).No previous algorithm for simultaneous gate and wire sizing can guarantee met solutions for general circuits.Our algorithm is an iterative one with a guarantee on convergence to global Oph.malsolutions.It is based on hgransian relmtion and "one-gatdwire-at-a-time" local optimi~tions, and is mtremely economical and fast.For example, we can optimize a circuit with 27,648 gates and wires in about 36 minutes using under 23 MB metnory on an IBM RS/6000 worhtation. Charlie Chung-Ping Chen, Chris C. N. Chu, Martin D. F. Wong |
ICCAD | 1 |
| 1997 | Optimal Wire-Sizing Function with Fringing Capacitance ConsiderationabstractIn this paper, we consider non-uniform wire-sizing under theElmore delay model.Given a wire segment of length L, letf(x) be the width of the wire at position x, 0 ≤ x ≤ L.It was shown in [Optimal Wire-sizing formula under the Elmore delay model, Shaping a distributed-RC line to minimize Elmore delay] that the optimal wire-sizing functionwhich minimizes delay is an exponential tapering functionf(x) = ae{-bx}, where a > 0 and b > 0 are constants.Unfortunately, [Optimal Wire-sizing formula under the Elmore delay model, Shaping a distributed-RC line to minimize Elmore delay] did not consider fringing capacitancewhich is at least comparable in size to area capacitance indeep submicron designs.As a result, exponential taperingis no longer the optimal strategy.In this paper, we showthat the optimal wire-sizing function, taking fringing capacitanceinto consideration, is f(x) = \frac{{ - c_f }}{{2c_0 }}(\frac{1}{{W(\frac{{ - c_f }}{{ae^{ - bx} }})}} + 1) whereW(x) = \sum\nolimits_{n = 1}^\infty{\frac{{( - n)^{n - 1} }}{{n!}}} x^n is the Lambert's W function, c{f}and c{0} are the respective fringing capacitance and area capacitanceof wire per unit square, a > 0 and b> 0 are constants.The optimal wire-sizing function degenerates into an exponentialtapering function as c}{f} = 0, and degenerates into asquare-root tapering function (f(x)=\sqrt {b - ax}, where a > 0and b > 0) as c{f} → √.Our experimental results show thatthe optimal wire-sizing function can significantly reduce theinterconnection delay of exponentially tapered wires.In thecase where lower and upper bounds on the wire widths aregiven, the optimal wire-sizing function is a truncated versionof the above function.Finally, our optimal wire-sizing functioncan be iteratively applied to optimally size all the wiresegments in a routing tree for objectives such as minimizingweighted sink delay, minimizing maximum sink delay, orminimizing area subject to delay bounds at the sinks. Charlie Chung-Ping Chen, Martin D. F. Wong |
DAC | 1 |
| 1996 | Fast Performance-Driven Optimization for Buffered Clock Trees Based on Lagrangian RelaxationabstractDelay, power, skew, area, and sensitivity are the most important concerns in current clock-tree design.We present in this paper an algorithm for simultaneously optimizing the above objectives by sizing wires and buers in clock trees.Our algorithm, based o n L agrangian relaxation method, can optimally minimize delay, power, and area simultaneously with very low skew and sensitivity.With linear storage overall and linear runtime per iteration, our algorithm is extremely economical, fast, and accurate; for example, our algorithm can solve a 6201-wire-segment clock-tree p r oblem using about 1-minute runtime and 1.3-MB memory and still achieve pico-second precision on an IBM RS/6000 workstation. Charlie Chung-Ping Chen, Yao-Wen Chang, Martin D. F. Wong |
DAC | 1 |
| 1996 | Optimal Wire-Sizing Formular Under the Elmore Delay ModelabstractIn this paper, we consider non-uniform wire-sizing.Given a wire s e gment of length L, let f(x) be the width of the wire at position x, 0 x L .We show that the optimal wiresizing function that minimizes the Elmore delay through the wire i s f ( x ) = ae bx , where a > 0 and b > 0 are c onstants that can be c omputed i n O (1) time.In the case where lower bound (L > 0 ) and upper bound (U > 0 ) on the wire widths are given, we show that the optimal wire-sizing function f(x) is a truncated version of ae bx that can also be determined i n O (1) time.Our wire-sizing formula can be iteratively applied to optimally size the wire s e gments in a routing tree. Charlie Chung-Ping Chen, Yao-Ping Chen, Martin D. F. Wong |
DAC | 1 |
| 1996 | Optimal non-uniform wire-sizing under the Elmore delay modelabstractWe consider non-uniform wire-sizing for general routing trees under the Elmore delay model. Three minimization objectives are studied: (1) total weighted sink-delays; (2) total area subject to sink-delay bounds; and (3) maximum sink delay. We first present an algorithm NWSA-wd for minimizing total weighted sink-delays based on iteratively applying the wire-sizing formula in [1]. We show that NWSA-wd always converges to an optimal wire-sizing solution. Based on NWSA-wd and the Lagrangian relaxation technique, we obtained two algorithms NWSA-db and NWSA-md which can optimally solve the other two minimization objectives. Experimental results show that our algorithms are efficient both in terms of runtime and storage. For example, NWSA-wd, with linear runtime and storage, can solve a 6201-wire segment routing-tree problem using about 1.5-second runtime and 1.3-MB memory on an IBM RS/6000 workstation. Charlie Chung-Ping Chen, Hai Zhou 0001, Martin D. F. Wong |
ICCAD | 1 |