EDBT 2026 Demo / reviewers in the wild / expert
Shih-Chieh Chang 0001
dblp:57/2331-1
· DBLP profile ↗
130ranked-venue papers
20as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 120 · 20 first-author · 9 since 2021Software engineering, systems software and programming languages · 11 · 3 since 2021Computer networks · 4Artificial intelligence and machine learning · 3Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Real-Time Dynamic IR-drop Prediction for IR ECOabstractDuring the IR Engineering Change Order (ECO) stage, cell moving leads to uncertain IR-drop results, requiring designers to explore multiple ECO candidates in each iteration to find a solution that effectively mitigates IR-drop, resulting in a long evaluation time. Although machine learning (ML)-based predictors have been proposed to expedite IR-drop evaluation, partial simulations are still needed to update features after ECO, taking over an hour and delaying IR-drop results. In this work, we propose a real-time dynamic IR-drop estimation method based on an XGBoost model with a global view of a cell’s surroundings. After ECO, our method provides dynamic IR-drop results in minutes without running any simulations and thus achieves real-time estimation. This allows designers to evaluate multiple ECO candidates concurrently in a single iteration. We conducted the experiments on five ECO candidates of an industrial design with 3 nm technology. The results show that the proposed model can effectively predict the IR-drop variations of moved cells after ECO with over $93 \%$ of fixed cells detected and an average MAE of 8.75 mV achieved. Furthermore, our method achieves an $88 X$ speedup over Voltus (commercial tool) and a $64 X$ speedup over traditional ML predictors when evaluating a single ECO candidate. The speedup is expected to increase as the number of ECO candidates increases. Yu-Che Lee, Yu-Chen Cheng, Yong-Fong Chang, Jia-Wei Lin, Hsun-Wei Pao, Yung-Chih Chen, Yi-Ting Li, Wuqian Tang, Shih-Chieh Chang 0001, Chun-Yao Wang |
DAC | 10 |
| 2025 | Dynamic IR-Drop Prediction Through a Multi-Task U-Net with Package Effect ConsiderationabstractDynamic IR drop analysis is a critical step in the design signoff stage for verifying the power integrity of a chip. Since the analysis is extremely time-consuming, it has led to the emergence of machine learning (ML)-based methods to expedite the procedure. While previous ML approaches have demonstrated the feasibility of IR drop prediction, they often neglect package effects and do not address diverse IR criteria for memory and standard cells. Thus, this paper introduces a novel ML-based approach designed for a fast and accurate prediction of multi-type IR drop, considering package effects. We develop new package-related features to account for the package impact on IR drop. The proposed model is based on a multitask U-net architecture that not only predicts two types of IR drops simultaneously but also increases prediction accuracy through comprehensive learning. To further enhance the model performance, we introduce the Input Fusion Block (IFB), which unifies units across channels within the input feature maps, leading to improved prediction accuracy. The experimental results show the across-pattern transferability of the proposed IR drop prediction method, demonstrating an RMSE of less than SmV and an MAE of less than 2mV on the unseen simulation patterns. Additionally, our proposed method achieves a 5X speedup compared to the commercial tool. Yu-Chen Cheng, Yong-Fong Chang, Yu-Che Lee, Jia-Wei Lin, Hsun-Wei Pao, Hao-Yun Chen, Yung-Chih Chen, Chun-Yao Wang, Shih-Chieh Chang 0001 |
DATE | 12 |
| 2025 | CNN Model Optimization Using a Hybrid Approach of Genetic Algorithm-based Pruning and Retraining with Knowledge Distillation
Kuan-Ling Chou, Cheng-Lung Wang, Yung-Chih Chen, Wuqian Tang, Yi-Ting Li, Shih-Chieh Chang 0001, Chun-Yao Wang |
ACM Great Lakes Symposium on VLSI | 6 |
| 2024 | A Hybrid Approach to Reverse Engineering on Combinational CircuitsabstractReverse engineering is a process that converts low-level description to high-level one. In this paper, we propose a hybrid approach consisting of structural analysis and black-box testing to reverse engineering on combinational circuits. Our approach is able to convert combinational circuits from gate-level netlist to Register-Transfer Level (RT-level) design accurately and efficiently. We developed our approach and participated in Problem A of the 2022 CAD Contest @ ICCAD. The revised version of our program successfully converted most cases and achieved higher scores than the 1stplace team in the contest. Wuqian Tang, Yi-Ting Li, Kai-Po Hsu, Kuan-Ling Chou, You-Cheng Lin, Chia-Feng Chien, Tzu-Li Hsu, Yung-Chih Chen, Ting-Chi Wang, Shih-Chieh Chang 0001, TingTing Hwang, Chun-Yao Wang |
DATE | 10 |
| 2024 | IR drop Prediction Based on Machine Learning and Pattern ReductionabstractWith the advances in semiconductor technology, the sizes of transistors are getting smaller, which has led to an increasingly severe impact of IR drop. Consequently, this trend has amplified the significance of IR drop analysis within the realm of chip design. However, analyzing IR drop is resource-intensive and time-consuming, since numerous simulation patterns are required to verify the power integrity of circuits. Additionally, with every engineering change order (ECO) step, a reevaluation is necessary. In this paper, we propose a machine learning-based method to predict IR drop levels and present an algorithm for reducing simulation patterns, which could reduce the time and computing resources required for IR drop analysis within the ECO flow. Experimental results show that our approach can reduce the number of patterns by approximately 50%, thereby decreasing the analysis time while maintaining accuracy. Yong-Fong Chang, Yung-Chih Chen, Yu-Chen Cheng, Shu-Hong Lin, Che-Hsu Lin, Chun-Yuan Chen, Yu-Che Lee, Jia-Wei Lin, Hsun-Wei Pao, Shih-Chieh Chang 0001, Yi-Ting Li, Chun-Yao Wang |
ACM Great Lakes Symposium on VLSI | 11 |
| 2024 | CIMR-V: An End-to-End SRAM-based CIM Accelerator with RISC-V for AI Edge DeviceabstractComputing-in-memory (CIM) is renowned in deep learning due to its high energy efficiency resulting from highly parallel computing with minimal data movement. However, current SRAM-based CIM designs suffer from long latency for loading weight or feature maps from DRAM for large AI models. Moreover, previous SRAM-based CIM architectures lack end-to-end model inference. To address these issues, this paper proposes CIMR-V, an end-to-end CIM accelerator with RISC-V that incorporates CIM layer fusion, convolution/max pooling pipeline, and weight fusion, resulting in an 85.14% reduction in latency for the keyword spotting model. Furthermore, the proposed CIM-type instructions facilitate end-to-end AI model inference and full stack flow, effectively synergizing the high energy efficiency of CIM and the high programmability of RISC-V. Implemented using TSMC 28nm technology, the proposed design achieves an energy efficiency of 3707.84 TOPS/W and 26.21 TOPS at 50 MHz. Yan-Cheng Guo, Tian-Sheuan Chang, Chih-Sheng Lin, Bo-Cheng Chiou, Chih-Ming Lai, Shyh-Shyuan Sheu, Wei-Chung Lo, Shih-Chieh Chang 0001 |
ISCAS | 8 |
| 2023 | Optimization of AI SoC with Compiler-assisted Virtual Design PlatformabstractAs deep learning keeps evolving dramatically with rapidly increasing complexity, the demand for efficient hardware accelerators has become vital. However, the lack of software/hardware co-development toolchains makes designing AI SoCs (artificial intelligent system-on-chips) considerably challenging. This paper presents a compiler-assisted virtual platform to facilitate the development of AI SoCs from the early design stage. The electronic system-level design platform provides rapid functional verification and performance/energy analysis. Cooperating with the neural network compiler, AI software and hardware can be co-optimized on the proposed virtual design platform. Our Deep Inference Processor is also utilized on the virtual design platform to demonstrate the effectiveness of the architectural evaluation and exploration methodology. Chih-Tsun Huang, Juin-Ming Lu, Yao-Hua Chen, Ming-Chih Tung, Shih-Chieh Chang 0001 |
ISPD | 5 |
| 2022 | Robust Binary Neural Network against Noisy Analog ComputationabstractComputing in memory (CIM) technology has shown promising results in reducing the energy consumption of a battery-powered device. On the other hand, to reduce MAC operations, Binary neural networks (BNN) show the potential to catch up with a full-precision model. This paper proposes a robust BNN model applied to the CIM framework, which can tolerate analog noises. These analog noises caused by various variations, such as process variation, can lead to low inference accuracy. We first observe that the traditional batch normalization can cause a BNN model to be susceptible to analog noise. We then propose a new approach to replace the batch normalization while maintaining the advantages. Secondly, in BNN, since noises can be removed when inputs are zeros during the multiplication and accumulation (MAC) operation, we also propose novel methods to increase the number of zeros in a convolution output. We apply our new BNN model in the keyword spotting application. Our results are very exciting. Zong-Han Lee, Fu-Cheng Tsai, Shih-Chieh Chang 0001 |
DATE | 3 |
| 2021 | Dynamic Workload Allocation for Edge ComputingabstractArtificial intelligence models implemented in power-efficient Internet-of-Things (IoT) devices have accuracy degradation due to limited power consumption. To mitigate the accuracy loss on IoT devices, an edge-server joint inference system is introduced. On the edge-server inference system, allocate more workloads to the server end can mitigate accuracy loss, but data transmission contributes to the power consumption of the edge device. Thus, in this article, we present a novel two-stage method to allocate workloads to the server or the edge to maximize inference accuracy under a power constraint. In the first stage, we present a clusterwise threshold-based method for estimating the trustworthiness of a prediction made at the edge. In the second stage, we further determine the workload allocation of a trustworthy image based on the probability of the top 1 prediction and the power constraint. In addition, we propose a fine-tuning process to the pretrained model at the edge for achieving better accuracy. In the experiments, we apply the proposed method to several well-known deep neural network models. The results show that the proposed method can improve inference accuracy up to 3.93% under a specific power constraint compared to previous methods. Yi-Wen Hung, Yung-Chih Chen, Chi Lo, Austin Go So, Shih-Chieh Chang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2020 | Person Identification by Walking Gesture Using Skeleton Sequences
Chu-Chien Wei, Li-Huang Tsai, Hsin-Ping Chou, Shih-Chieh Chang 0001 |
ACIVS | 4 |
| 2020 | Accuracy Tolerant Neural Networks Under Aggressive Power OptimizationabstractWith the success of deep learning, many neural network models have been proposed and applied to various applications. In several applications, the devices used to implement the complicated models have limited power resources, and thus aggressive optimization techniques are often applied for saving power. However, some optimization techniques, such as voltage scaling and multiple threshold voltages, may increase the probability of error occurrence due to slow signal propagation, which increases the path delay in a circuit and fails some input patterns. Although neural network models are considered to have some error tolerance, the prediction accuracy could be significantly affected, when there are a large number of errors. Thus, in this paper, we propose a scheme to mitigate the errors caused by slow signal propagation. Since the delay of multipliers dominates the critical path, we consider the patterns significantly altered by the slow signal propagation in a multiplier. We propose two methods, weight distribution and error-aware quantization to prevent the patterns from failure. Since we modify a neural network on the software side and it is unnecessary to re-design the hardware structure. The experimental results show that the proposed scheme is effective for several neural network models. It can improve the network accuracy by up to 27% under the consideration of slow signal propagation. Xiang-Xiu Wu, Yi-Wen Hung, Yung-Chih Chen, Shih-Chieh Chang 0001 |
DATE | 4 |
| 2019 | Aging-aware chip health prediction adopting an innovative monitoring strategyabstractConcerns exist that the reliability of chips is worsening because of downscaling technology. Among various reliability challenges, device aging is a dominant concern because it degrades circuit performance over time. Traditionally, runtime monitoring approaches are proposed to estimate aging effects. However, such techniques tend to predict and monitor delay degradation status for circuit mitigation measures rather than the health condition of the chip. In this paper, we propose an aging-aware chip health prediction methodology that adapts to workload conditions and process, supply voltage, and temperature variations. Our prediction methodology adopts an innovative on-chip delay monitoring strategy by tracing representative aging-aware delay behavior. The delay behavior is then fed into a machine learning engine to predict the age of the tested chips. Experimental results indicate that our strategy can obtain 97.40% accuracy with 4.14% area overhead on average. To the authors' knowledge, this is the first method that accurately predicts current chip age and provides information regarding future chip health. Yun-Ting Wang, Kai-Chiang Wu, Chung-Han Chou, Shih-Chieh Chang 0001 |
ASP-DAC | 4 |
| 2019 | Improving Adversarial Robustness via Guided Complement EntropyabstractAdversarial robustness has emerged as an important topic in deep learning as carefully crafted attack samples can significantly disturb the performance of a model. Many recent methods have proposed to improve adversarial robustness by utilizing adversarial training or model distillation, which adds additional procedures to model training. In this paper, we propose a new training paradigm called Guided Complement Entropy (GCE) that is capable of achieving "adversarial defense for free," which involves no additional procedures in the process of improving adversarial robustness. In addition to maximizing model probabilities on the ground-truth class like cross-entropy, we neutralize its probabilities on the incorrect classes along with a "guided" term to balance between these two terms. We show in the experiments that our method achieves better model robustness with even better performance compared to the commonly used cross-entropy training objective. We also show that our method can be used orthogonal to adversarial training across well-known methods with noticeable robustness gain. To the best of our knowledge, our approach is the first one that improves model robustness without compromising performance. Hao-Yun Chen, Jhao-Hong Liang, Shih-Chieh Chang 0001, Jia-Yu Pan, Yuting Chen 0002, Wei Wei 0019, Da-Cheng Juan |
ICCV | 3 |
| 2019 | Complement Objective Training
Hao-Yun Chen, Pei-Hsin Wang, Chun-Hao Liu, Shih-Chieh Chang 0001, Jia-Yu Pan, Yuting Chen 0002, Wei Wei 0019, Da-Cheng Juan |
ICLR (Poster) | 4 |
| 2019 | Hierarchical LSTM: Modeling Temporal Dynamics and Taxonomy in Location-Based Mobile Check-Ins
Chun-Hao Liu, Da-Cheng Juan, Xuan-An Tseng, Wei Wei 0019, Yuting Chen 0002, Jia-Yu Pan, Shih-Chieh Chang 0001 |
PAKDD (2) | 7 |
| 2018 | Searching toward pareto-optimal device-aware neural architecturesabstractRecent breakthroughs in Neural Architectural Search (NAS) have achieved state-of-the-art performance in many tasks such as image classification and language understanding. However, most existing works only optimize for model accuracy and largely ignore other important factors imposed by the underlying hardware and devices, such as latency and energy, when making inference. In this paper, we first introduce the problem of NAS and provide a survey on recent works. Then we deep dive into two recent advancements on extending NAS into multiple-objective frameworks: MONAS [19] and DPP-Net [10]. Both MONAS and DPP-Net are capable of optimizing accuracy and other objectives imposed by devices, searching for neural architectures that can be best deployed on a wide spectrum of devices: from embedded systems and mobile devices to workstations. Experimental results are poised to show that architectures found by MONAS and DPP-Net achieves Pareto optimality w.r.t the given objectives for various devices. An-Chieh Cheng, Jin-Dong Dong, Chi-Hung Hsu, Shu-Huan Chang, Min Sun 0001, Shih-Chieh Chang 0001, Jia-Yu Pan, Yuting Chen 0002, Wei Wei 0019, Da-Cheng Juan |
ICCAD | 6 |
| 2018 | Sensor-Based Time Speculation in the Presence of Timing VariabilityabstractTime speculation has been widely used to achieve high performance in modern design as it exploits average-case timing optimization instead of worst-case timing optimization focusing on reducing longest path delay which rarely happens. Variable-latency design (VLD) style is one research category of time speculation. Since process and environmental variations are hard to predict, traditional variable-latency units (VLUs) designed at presilicon stage will suffer significant performance loss due to pessimistic assumptions for addressing variations. In this paper, we propose a novel sensor-based, transition-aware VLU (S-VLU) scheme adapting to process-voltage-temperature (PVT) variations by using in situ sensors to obtain real-time transition information in a circuit. Moreover, we also propose a sensor deployment strategy to achieve near-maximal performance gain. On average, the S-VLU achieves a 31.27% performance improvement as compared to a 19.26% improvement by using traditional HL. The area overhead of the S-VLU is 13.48%. To the best of the authors' knowledge, this is the first wok to address PVT variations in VLD style. Chung-Han Chou, Tsui-Yun Chang, Kai-Chiang Wu, Shih-Chieh Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Contactless Testing for Prebond Interposers
Kai-Hsiang Hsu, Yung-Chih Chen, You-Luen Lee, Shih-Chieh Chang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | An Adaptive Mechanism for Designing Efficient Snoop Filters
Sze-Chen Cho, Shih-Chieh Chang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Pattern based runtime voltage emergency prediction: An instruction-aware block sparse compressed sensing approachabstractThe relentless technology scaling calls for reduced supply voltage for dynamic power suppression. On the other hand, transistor threshold voltage cannot be scaled at the same pace to avoid excessive leakage power. Consequently, the noise margin is significantly reduced, leading to the deployment of various noise management systems that handle runtime voltage emergencies. Most of these systems rely on on-chip noise sensors, which are large in size and consume significant power. To tackle this issue, in this paper we propose a sensor-less voltage emergency estimation framework. It explores the relationship between switching activities and noise, and takes advantage of block sparse compressed sensing developed by the signal processing society. Experimental results on a few industrial designs show that by monitoring registers, voltage emergencies can be successfully predicted. Yu-Guang Chen, Michihiro Shintani, Takashi Sato 0001, Yiyu Shi 0001, Shih-Chieh Chang 0001 |
ASP-DAC | 5 |
| 2017 | CN-SIM: A cycle-accurate full system power delivery noise simulatorabstractThis paper introduces CN-SIM, a cycle accurate, full system, power delivery (PD) noise simulator. CN-SIM provides a cross layer connectivity form application layer, to the architecture layer, to the circuit layer, which is much needed to realistically estimate PD noise. Thus, making it easier for system architects to explore multilayer design optimizations. CN-SIM's granularity at its deepest is at the functional unit (FU) level. The experimental results of running PARSEC suite benchmarks for different system configurations and different industrial PD design have illustrated CN-SIM's capability to capture the crosslayer impact on PD noise. Kassan Unda, Chung-Han Chou, Shih-Chieh Chang 0001, Cheng Zhuo, Yiyu Shi 0001 |
ASP-DAC | 3 |
| 2017 | A Dynamic Deep Neural Network Design for Efficient Workload Allocation in Edge ComputingabstractUnreliable communication channels and limited computing resources at the edge end are two primary constraints of battery-powered movable devices, such as autonomous robots and unmanned aerial vehicles (UAVs). The impact is especially severe for those performing deep neural network (DNN) computations. With increasing demand for accuracy, the trend in modern DNN designs is the use of cascaded modularized layers. Implementing a deep network at the edge increases computational workloads and resource occupancy, leading to an increase in battery drain. Using a shallow network and offloading workloads to backbone servers, however, incur significant latency overheads caused by unstable communication channels. Hence, dynamic DNN design techniques for efficient workload allocation are urgently required to manage the amount of workload transmissions while achieving the required accuracy. In this paper, we explore the use of authentic operation (AO) unit and dynamic network structure to enhance DNNs. The AO unit defines a set of stochastic threshold values for different DNN output classes and determines at runtime if an input has to be transferred to backbone servers for further analysis. The dynamic network structure adjusts its depth according to channel availability. Experiments have been comprehensively performed on several well-known DNN models and datasets. Our results show that, on an average, the proposed techniques are able to reduce the amount of transmissions by up to 17% compared to previous methods under the same accuracy requirement. Chi Lo, Yu-Yi Su, Chun-Yi Lee, Shih-Chieh Chang 0001 |
ICCD | 4 |
| 2017 | DC-Prophet: Predicting Catastrophic Machine Failures in DataCenters
You-Luen Lee, Da-Cheng Juan, Xuan-An Tseng, Yuting Chen 0002, Shih-Chieh Chang 0001 |
ECML/PKDD (3) | 5 |
| 2017 | Ping-Pong Mesh: A New Resonant Clock Design for Surge Current and Area Overhead ReductionabstractIn advanced technologies, on-chip-variation (OCV) has accounted for a large proportion of clock skew, which limits the performance of a circuit. To mitigate the OCV problem, a mesh structure has been widely used in high-performance designs. Unfortunately, clock mesh structure also causes large power consumption and large power-ground surge current. Therefore, recently, several approaches have been proposed to apply resonant clock to reduce power consumption. However, previous works often suffer from area overhead because of the need to insert large decoupling capacitors. In this paper, we propose a novel resonant clock mesh structure, called ping-pong mesh, to overcome these drawbacks. Our ping-pong mesh contains two submeshes, each of which plays the role of the decoupling capacitor of the other, and the clocks in two submeshes operate in completely opposite phases. Our ping-pong mesh has the following two advantages: 1) a ping-pong mesh does not need additional decoupling capacitors as in previous works and 2) a ping-pong mesh can reduce the power-ground surge current about half of previous works. Benchmark data consistently show that our ping-pong mesh does work well in practice. Chung-Han Chou, Yenting Lai, Yi-Chun Chang, Chih-Yu Wang 0002, Liang-Chia Cheng, Shih-Hsu Huang, Shih-Chieh Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2017 | Leak Stopper: An Actively Revitalized Snoop Filter Architecture with Effective Generation ControlabstractTo alleviate high energy dissipation of unnecessary snooping accesses, snoop filters have been designed to reduce snoop lookups. These filters have the problem of decreasing filtering efficiency, and thus usually rely on partial or whole filter reset by detecting block evictions. Unfortunately, the reset conditions occur infrequently or unevenly (called passive filter deletion ). This work proposes the concept of revitalized snoop filter (RSF) design, which can actively renew the destination filter by employing a generation wrapping-around scheme for various reference behaviors. We further utilize a sampling mechanism for RSF to timely trigger precise filter revitalizations, so that unnecessary RSF flushing can be minimized. The proposed RSF can be integrated to various existent inclusive snoop filters with only a minor change to their designs. We evaluate our proposed design and demonstrate that RSF eliminates 58.6% of snoop energy compared to JETTY on average while inducing only 6.5% of revitalization energy overhead. In addition, RSF eliminates 45.5% of snoop energy compared to stream registers on average and only induces 2.5% of revitalization energy overhead. Overall, these RSFs reduce the total L2 cache energy consumption by 52.1% (58.6% -- 6.5%) as compared to JETTY and by 43% (45.5% -- 2.5%) as compared to stream registers. Furthermore, RSF improves the overall performance by 1% to 1.4% on average compared to JETTY and stream registers for various benchmark suites. Yin-Chi Peng, Chien-Chih Chen, Hsiang-Jen Tsai, Keng-Hao Yang, Pei-Zhe Huang, Shih-Chieh Chang 0001, Wen-Ben Jone, Tien-Fu Chen |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2017 | Perfect Hashing Based Parallel Algorithms for Multiple String Matching on Graphic Processing UnitsabstractMultiple string matching has a wide range of applications such as network intrusion detection systems, spam filters, information retrieval systems, and bioinformatics. To accelerate multiple string matching, many hardware approaches are proposed to accelerate string matching. Among the hardware approaches, memory architectures have been widely adopted because of their flexibility and scalability. A conventional memory architecture compiles multiple string patterns into a state machine and performs string matching by traversing the corresponding state transition table. Due to the ever-increasing number of attack patterns, the memory used for storing the state transition table increased tremendously. Therefore, memory reduction has become a crucial issue in optimizing memory architectures. In this paper, we propose two parallel string matching algorithms which adopt perfect hashing to compact a state transition table. Different from most state-of-the-art approaches implemented on specific hardware such as TCAM, FPGA, or ASIC, our proposed approaches are easily implemented on commodity DRAM and extremely suitable to be implemented on GPUs. The proposed algorithms reduce up to 99.5 percent memory requirements for storing the state transition table compared to the traditional two-dimensional memory architecture. By studying existing approaches, our results obtain significant improvements in memory efficiency. Jin-Cheng Li, Chen-Hsiung Liu, Shih-Chieh Chang 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | EditorialabstractAs I start my second two-year term (2017–2018) as the Editor-in-Chief (EIC) of the IEEE Transactions on Very Large Scale Integration Systems (TVLSI), I wish the TVLSI readership a very happy new year and continued professional success. It gives me great pleasure to report on the state of the journal and our performance metrics. Over the past two years, TVLSI has seen a healthy increase in the number of submissions—from 687 in 2014 to 770 in 2015, and at the time of writing of this editorial, we are at 760 submissions for 2016. We expect the number of submissions for 2016 to cross 800 before the end of the year. TVLSI, therefore, continues to be the premier archival journal for university researchers and industry practitioners in the broad area of VLSI system design. Krishnendu Chakrabarty, Massimo Alioto, Bevan M. Baas, Chirn Chye Boon, Meng-Fan Chang, Naehyuck Chang, Yao-Wen Chang, Chip-Hong Chang, Shih-Chieh Chang 0001, Poki Chen, Masud H. Chowdhury, Pasquale Corsonello, Ibrahim M. Elfadel, Said Hamdioui, Masanori Hashimoto, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Rajiv V. Joshi, Tanay Karnik, Mehran Mozaffari Kermani, Chulwoo Kim, Jaydeep P. Kulkarni, Eren Kursun, Erik Larsson, Hai Li 0001, Huawei Li 0001, Patrick P. Mercier, Prabhat Mishra 0001, Makoto Nagata, Arun Natarajan 0001, Koji Nii, Partha Pratim Pande, Ioannis Savidis, Mingoo Seok, Sheldon X.-D. Tan, Mark Tehranipoor, Aida Todri, Miroslav N. Velev, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2016 | A novel low-cost dynamic logic reconfigurable structure strategy for low power optimizationabstractLow power design techniques have been extensively applied in modern IC designs to avoid negative side effects from high power density. Unlike Dynamic Voltage and/or Frequency Scaling (DVFS) approaches only applied on a “fixed” design, we propose a dynamic logic reconfigurable structure strategy which allows dynamic switching from a high speed/power logic structure to a low speed/power logic structure. A design with such configurable structure is called Dynamic Logic Reconfigurable Structure (DLRS). Different from approximate computing which trades off between computation accuracy and power, our DLRS designs maintain data integrity. In this paper, we propose novel low-cost DLRS adders and multipliers, and a comprehensive framework for low power designs. We further integrate DLRS with DVFS, which creates more flexibility to trade-off between performance and power consumption. Experimental results show that with DLRS adders and multipliers in three indoor designs, the proposed method can achieve up to 60.05% power reduction compared with traditional DVFS scheme with only 6.55% area overhead. Yu-Guang Chen, Wan-Yu Wen, Yun-Ting Wang, You-Luen Lee, Shih-Chieh Chang 0001 |
ASP-DAC | 5 |
| 2016 | High-Performance Deadlock-Free ID Assignment for Advanced Interconnect ProtocolsabstractIn a modern system-on-chip design, hundreds of cores and intellectual properties can be integrated into a single chip. To be suitable for high-performance interconnects, designers increasingly adopt advanced interconnect protocols that support novel mechanisms of parallel accessing, including outstanding transactions and out-of-order completion of transactions. To implement those novel mechanisms, a master tags an ID to each transaction to decide in-order or out-of-order properties. However, these advanced protocols may lead to transaction deadlocks that do not occur in traditional protocols. To prevent the deadlock problem, current solutions stall suspicious transactions and in certain cases, many such stalls can incur serious performance penalty. In this brief, we propose a novel ID assignment mechanism that guarantees the issued transactions to be deadlock-free and results in significant reduction in the number of transaction stalls issued by masters. Our experimental results show encouraging performance improvements compared with previous works with little hardware and power overheads. Hsuan-Ming Chou, Yi-Chiao Chen, Keng-Hao Yang, Jean Tsao, Shih-Chieh Chang 0001, Wen-Ben Jone, Tien-Fu Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2016 | Skew Minimization With Low Power for Wide-Voltage-Range Multipower-Mode DesignsabstractIn a multipower-mode design, as the range of the supply voltage becomes wide, a large clock skew may occur among different power domains. To remove this clock skew, conventional power-mode-aware buffers (PMABs) require a large overhead on power consumption. In this brief, we propose a new PMAB architecture for wide-voltage-range multipower-mode designs. The proposed PMAB architecture is composed of two serially connected sub-PMABs at two different voltage levels, respectively. In the front sub-PMAB, the low voltage level is used for coarse-grained clock skew minimization. In the back sub-PMAB, the high voltage level is used for fine-grained clock skew minimization. Benchmark data show that the proposed approach can effectively eliminate the clock skew with small power consumption. Chung-Han Chou, Hua-Hsin Yeh, Shih-Hsu Huang, Yow-Tyng Nieh, Shih-Chieh Chang 0001, Yung-Tai Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2015 | Hybrid coverage assertions for efficient coverage analysis across simulation and emulation environmentsabstractCoverage metrics are commonly used to measure the completeness of the verification test suites. However, in a modern hardware-accelerated environment, coverage may be analyzed across a simulator and an emulator. Hence, neither conventional coverage techniques performed in simulation nor hardware coverage monitors embedded in an emulator can be directly applied. To resolve the above problem, we propose using coverage assertions to detect coverage events across a simulator and an emulator. In addition, an Assertion Operation Graph and graph-based algorithms are proposed to minimize the hardware and performance overheads of coverage assertions. We perform experiments in the hardware-accelerated environment of Xilinx ISE and show an encouraging reduction of coverage assertion overheads. Hsuan-Ming Chou, Hong-Chang Wu, Yi-Chiao Chen, Jean Tsao, Shih-Chieh Chang 0001 |
ASP-DAC | 5 |
| 2015 | Q-Learning Based Dynamic Voltage Scaling for Designs with Graceful DegradationabstractDynamic voltage scaling (DVS) has been widely used to suppress power consumption in modern designs. The decision of optimal operating voltage at runtime should consider the variations in workload, process as well as environment. As these variations are hard to predict accurately at design time, various reinforcement learning based DVS schemes have been proposed in the literature. However, none of them can be readily applied to designs with graceful degradation, where timing errors are allowed with bounded probability to trade for further power reduction. In this paper, we propose a Q-learning based DVS scheme dedicated to the designs with graceful degradation. We compare it with two deterministic DVS schemes, i.e., a stepping based scheme and a statistical modeling based scheme. Experimental results on three 45nm industrial designs show that the proposed Q-learning based scheme can achieve up to 83.9% and 29.1% power reduction respectively with 0.01 timing error probability bound. To the best of the authors' knowledge, this is the first in-depth work to explore reinforcement learning based DVS schemes for designs with graceful degradation. Yu-Guang Chen, Wan-Yu Wen, Yiyu Shi 0001, Shih-Chieh Chang 0001 |
ISPD | 5 |
| 2015 | Novel Spare TSV Deployment for 3-D ICs Considering Yield and Timing ConstraintsabstractIn 3-D integrated circuits, through silicon via (TSV) is a critical enabling technique to provide vertical connections. However, it may suffer from many reliability issues such as undercut, misalignment, or random open defects. Various fault-tolerance mechanisms have been proposed in literature to improve yield, at the cost of significant area overhead. In this paper, we focus on the structure that uses one spare TSV for a group of original TSVs, and study the optimal assignment of spare TSVs under yield and timing constraints to minimize the total area overhead. We show that such problem can be modeled as a constrained graph decomposition problem. Two efficient heuristics are further developed to address this problem. Experimental results show that under the same yield and timing constraints, our heuristic can reduce the area overhead induced by the fault-tolerance mechanisms by up to 61%, compared with a seemingly more intuitive nearest-neighbor-based heuristic. Yu-Guang Chen, Wan-Yu Wen, Yiyu Shi 0001, Wing-Kai Hon, Shih-Chieh Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2015 | Soft-Error-Tolerant Design Methodology for Balancing Performance, Power, and ReliabilityabstractSoft error has become an important reliability issue in advanced technologies. To tolerate soft errors, solutions suggested in previous works incur significant performance and power penalties, especially when a design with fault-tolerant structures is overprotected. In this paper, we present a soft-error-tolerant design methodology to tradeoff performance, power, and reliability for different applications. First, four novel detection and correction flip-flop (FF) structures are proposed to provide different levels of tolerance capability against soft errors. Second, architecture-level vulnerability and logic-level susceptibility analyses are employed to identify weak FFs that can easily cause program execution errors. Third, an optimization framework is developed to synthesize the proposed four novel FF structures into weak and highly observable storage bits with the flexibility of trading off performance, power, and reliability. A five-stage pipeline RISC core (UniRISC) is adopted to demonstrate the usefulness of our methodology. Experimental results show that the proposed method can accomplish design goals by balancing performance, power, and reliability. For example, we can not only satisfy the reliability requirement that no more than five errors occur per one billion hours in a design but also reduce up to 87% performance overhead and 91% power overhead when compared with previous works. Hsuan-Ming Chou, Ming-Yi Hsiao, Yi-Chiao Chen, Keng-Hao Yang, Jean Tsao, Chiao-Ling Lung, Shih-Chieh Chang 0001, Wen-Ben Jone, Tien-Fu Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2015 | A Fault Detection and Tolerance Architecture for Post-Silicon Skew TuningabstractClock skew minimization that is an important issue in very large scale integration design has become difficult due to the presence of process, voltage, and temperature (PVT) variations. The post-silicon skew tuning (PST) technique with the ability to tolerate PVT variations, even after a chip is manufactured has generated considerable discussion. The basic idea of the PST architecture is to minimize the clock skew dynamically. Unlike most previous works that have focused on the implementation and the performance issues of a PST architecture, this paper focuses on the testing issues of a PST architecture. However, testing the variation tolerance ability of the PST architecture is difficult because the clock skew does not directly affect the functionality of a design. In this paper, we propose an efficient fault model considering the physical limitation of the devices for the PST architecture. In addition, we propose some novel structures to detect the manufacturing faults and increase the robustness of a PST architecture. Our experiment shows that with a little increase in overhead, we can achieve robustness. Mac Y. C. Kao, Kun-Ting Tsai, Shih-Chieh Chang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Critical Path Monitor Enabled Dynamic Voltage Scaling for Graceful Degradation in Sub-Threshold DesignsabstractSub-threshold designs play an important role in energy-constrained applications. In those designs, path delays depend exponentially on threshold voltage/temperature. As such, dynamic configurations at runtime are desired for best trade-off between operating power and performance. Unfortunately, most existing works only consider either process or temperature variations but not both, resulting in sub-optimal configurations or even functional failures. Moreover, little study has been performed on the graceful degradation of sub-threshold designs, which is important in the presence of drastic delay variations. Towards this, we present a novel critical path monitor based dynamic voltage scaling scheme. Considering both process and temperature variations, it minimizes the operating power under a given timing error probability (TEP) bound. An exact method to decide the optimal switching thresholds is also proposed. Experimental results on 45nm industrial designs show that with only 1% TEP, our scheme can reduce the operating power by up to 75.3% compared with the constant voltage scheme. To the best of the authors' knowledge, this is the very first work on dynamic configuration for graceful degradation in sub-threshold designs. Yu-Guang Chen, Kuan-Yu Lai, Wan-Yu Wen, Yiyu Shi 0001, Shih-Chieh Chang 0001 |
DAC | 6 |
| 2014 | Contactless Stacked-die Testing for Pre-bond InterposersabstractA stacked-die product integrates multiple dies on interposers. In this paper, we first discuss the difficulties of traditional testing mechanism for interposers. To improve production yield, a contactless testing mechanism for pre-bond interposers is proposed. Our testing mechanism attempts to detect a defective interposer from the thermal image after heating the interposer. We propose to extract special features from the thermal image and then use a clustering algorithm to determine whether the interposer is defective. Experimental results show that our testing mechanism can efficiently improve the yield from 70.5% to 96.84%. Jui-Hung Chien, Ruei-Siang Hsu, Hsueh-Ju Lin, Ka-Yi Yeh, Shih-Chieh Chang 0001 |
DAC | 5 |
| 2014 | Yield and timing constrained spare TSV assignment for three-dimensional integrated circuitsabstractThrough Silicon Via (TSV) is a critical enabling technique in three-dimensional integrated circuits (3D ICs). However, it may suffer from many reliability issues. Various fault-tolerance mechanisms have been proposed in literature to improve yield, at the cost of significant area overhead. In this paper, we focus on the structure that uses one spare TSV for a group of original TSVs, and study the optimal assignment of spare TSVs under yield and timing constraints to minimize the total area overhead. We show that such problem can be modeled through constrained graph decomposition. An efficient heuristic is further developed to address this problem. Experimental results show that under the same yield and timing constraints, our heuristic can reduce the area overhead induced by the fault-tolerance mechanisms by up to 38%, compared with a seemingly more intuitive nearest-neighbor based heuristic. Yu-Guang Chen, Kuan-Yu Lai, Ming-Chao Lee, Yiyu Shi 0001, Wing-Kai Hon, Shih-Chieh Chang 0001 |
DATE | 6 |
| 2014 | Package geometric aware thermal analysis by infrared-radiation thermal imagesabstractSince packages affect the amount of heat transfer, it is important to include package and heat sink in thermal analysis. In this paper, we study the full-chip thermal response with different packages. We first discuss the difficulties of obtaining accurate package models for simulation. To facilitate a designer to perform thermal simulation with different packages, we propose to use a matrix called the package-transfer matrix which can transform a temperature profile of one package to another temperature profile of the desired package. To estimate and verify a package-transfer matrix, we propose an efficient method which uses Infrared Radiation (IR) images from two carefully design test chips with PBGA packages. Our experimental results show that the default package model CBGA in HotSpot can be accurately transferred to any other package through the package-transfer matrix. Jui-Hung Chien, Ruei-Siang Hsu, Hsueh-Ju Lin, Shih-Chieh Chang 0001 |
DATE | 5 |
| 2014 | Multibit Retention Registers for Power Gated Designs: Concept, Design, and DeploymentabstractRetention registers have been widely used in power gated designs to store data during sleep mode. However, their excessive area and leakage power render it imperative to minimize the total retention storage size. The current industry practice replaces all registers with singlebit retention ones, which significantly limits the design freedom and yields suboptimal designs. Toward this, for the first time in the literature, we propose the concept and the design of multibit retention registers, with which only selected registers need to be replaced. The technique can significantly reduce the number of bits that need to be stored and thus the leakage power, but needs several clock cycles for mode transition. In addition, an efficient assignment algorithm is developed to minimize the total retention storage size subject to mode transition latency constraint. Experimental results show that our framework on average can reduce the leakage power in sleep mode by 84% along with additional mode transition latency of 6 to 11 clock cycles, compared with the singlebit retention register-based design. Yu-Guang Chen, Hui Geng, Kuan-Yu Lai, Yiyu Shi 0001, Shih-Chieh Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2014 | A Fuzzy-Matching Model With Grid Reduction for Lithography Hotspot DetectionabstractIn advanced IC manufacturing, as the gap increases between lithography optical wavelength and feature size, it becomes challenging to detect problematic layout patterns called lithography hotspot. In this paper, we propose a novel fuzzy matching model which extracts appropriate feature vectors of hotspot and nonhotspot patterns. Our model can dynamically tune appropriate fuzzy regions around known hotspots. Based on this paper, we develop a fast algorithm for lithography hotspot detection with high accuracy of detection and low probability of false-alarm counts. In addition, since higher dimensional size of feature vectors can produce better accuracy but requires longer run time, this paper proposes a grid reduction technique to significantly reduce the CPU run time with very minor impact on the advantages of higher dimensional space. Our results are very encouraging, with average 94.5% accuracy and low false-alarm counts on a set of test benchmarks. Wan-Yu Wen, Jin-Cheng Li, Sheng-Yuan Lin, Jing-Yi Chen, Shih-Chieh Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2014 | Thermal-Aware On-Line Scheduler for 3-D Many-Core Processor Throughput Optimizationabstract3-D many-core processor (3-D MCP) has become an emerging technology to tackle the power wall problem due to rapidly increasing number of transistors. However, when maximizing the throughput of 3-D MCP, which is expressed as a weighted sum of the speeds, due to the inherent heat removal limitation, thermal issues must be taken into consideration. Since the temperature of a core strongly depends on its location in the 3-D IC, a proper task allocation can alleviate the thermal problem and improve the throughput. Nevertheless, conventional techniques require computationally intensive thermal simulation, which prohibits its usage from the online application. In this paper, we propose an efficient online task allocation and task migration algorithm attempting to maximize the throughput of 3-D MCP simultaneously, considering unfinished tasks left from the last scheduling interval and new incoming tasks of this scheduling interval. The results of our experiments show that our proposed method achieves a 20.82X runtime speedup. These results are comparable to the exhaustive solutions obtained from optimization-modeling software LINGO. In addition, on average, our throughput results, with and without consideration of unfinished tasks, are only 4.39% and 0.69% worse, respectively, than that of the exhaustive method. In 128 task-to-core allocations, our method takes only 0.951 ms, which is 59.39 times faster than that of the previous work. Cody Hao Yu, Chiao-Ling Lung, Yi-Lun Ho, Ruei-Siang Hsu, Ding-Ming Kwai, Shih-Chieh Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2013 | A novel fuzzy matching model for lithography hotspot detectionabstractIn advanced IC manufacturing, as the gap between lithography optical wavelength and feature size increases, it becomes challenging to detect problematic layout patterns called lithography hotspot. In this paper, we propose a novel fuzzy matching model which can dynamically tune appropriate fuzzy regions around known hotspots. Based on this model, we develop a fast algorithm for lithography hotspot detection with very low chances of false-alarm. Our results are very encouraging with under 0.56 CPU-hrs/mm2 runtime. Sheng-Yuan Lin, Jing-Yi Chen, Jin-Cheng Li, Wan-Yu Wen, Shih-Chieh Chang 0001 |
DAC | 5 |
| 2013 | Low-power timing closure methodology for ultra-low voltage designsabstractAs the supply voltage is down to the ultra-low voltage (ULV) level, timing closure becomes a serious challenge in the use of multiple power modes. Due to a wide voltage range, a very huge clock skew may occur among different power modes. To reduce this huge clock skew, the conventional power-mode-aware clock tree often suffers from a huge overhead on power consumption. Moreover, at the ULV level, since the setup time and the hold time of each register dramatically increase, the number of timing violations also increases greatly. However, the existing minimum padding technique cannot fix hold time violations in multiple power modes. Based on those two observations, in this paper, we propose a low-power timing closure methodology, which incorporates the synthesis of clock tree and data path, for multipower-mode ULV designs. Our low-power timing closure methodology has two main approaches. First, we use multiple power modes to build a power-mode-aware clock tree for reducing clock skew with very small power consumption. Second, we propose the first multi-power-mode minimum padding technique to fix all the hold time violations in all the power modes simultaneously. Experimental results consistently show that the integration of both approaches yields the best results. Wen-Pin Tu, Chung-Han Chou, Shih-Hsu Huang, Shih-Chieh Chang 0001, Yow-Tyng Nieh, Chien-Yung Chou |
ICCAD | 4 |
| 2013 | Accelerating Pattern Matching Using a Novel Parallel Algorithm on GPUsabstractGraphics processing units (GPUs) have attracted a lot of attention due to their cost-effective and enormous power for massive data parallel computing. In this paper, we propose a novel parallel algorithm for exact pattern matching on GPUs. A traditional exact pattern matching algorithm matches multiple patterns simultaneously by traversing a special state machine called an Aho-Corasick machine. Considering the particular parallel architecture of GPUs, in this paper, we first propose an efficient state machine on which we perform very efficient parallel algorithms. Also, several techniques are introduced to do optimization on GPUs, including reducing global memory transactions of input buffer, reducing latency of transition table lookup, eliminating output table accesses, avoiding bank-conflict of shared memory, coalescing writes to global memory, and enhancing data transmission via peripheral component interconnect express. We evaluate the performance of the proposed algorithm using attack patterns from Snort V2.8 and input streams from DEFCON. The experimental results show that the proposed algorithm performed on NVIDIA GPUs achieves up to 143.16-Gbps throughput, 14.74 times faster than the Aho-Corasick algorithm implemented on a 3.06-GHz quad-core CPU with the OpenMP. The library of the proposed algorithm is publically accessible through Google Code. Chen-Hsiung Liu, Lung-Sheng Chien, Shih-Chieh Chang 0001 |
IEEE Trans. Computers | 4 |
| 2013 | Through-Silicon Via Fault-Tolerant Clock Networks for 3-D ICsabstractClock network synthesis is one of the most important and challenging problems in 3-D ICs. The clock signals have to be delivered by through-silicon vias (TSVs) to different tiers with minimum skew. While there are a few related works in literature, none consider the reliability of TSVs in a clock tree. Accordingly, the failure of any TSV in the clock tree yields a bad chip. The naive solution using double-TSV can alleviate the problem, but the significant area overhead renders it less practical for large designs. In this paper, we propose a novel TSV fault-tolerant unit (TFU) to provide tolerance against TSV failures. The TFU makes use of the existing 2-D redundant trees designed for prebond testing, and thus has minimum area overhead. In addition, the number of TSVs in a TFU is also adjustable to allow flexibility during clock network synthesis. Compared with the conventional double TSV technique, the 3-D clock network constructed by TFUs can achieve 58% area overhead reduction with similar yield rate on an industrial case. To the best of the authors' knowledge, this is the first work in the literature that considers the fault tolerance of a 3-D clock network. It can be easily integrated with any bottom-up clock network synthesis algorithm. Chiao-Ling Lung, Yu-Shih Su, Hsih-Hsiu Huang, Yiyu Shi 0001, Shih-Chieh Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2013 | Embedding Repeaters in Silicon IPs for Cross-IP InterconnectionsabstractDuring systems-on-a-chip (SoC) integration, silicon intellectual properties (IPs) are generally regarded as blockages to long interconnections that connect different IPs. With this constraint, conventional designs are forced to place those repeaters that drive long interconnections outside the IP. These designs either lead to a longer interconnection distance requiring more repeaters or result in a longer signal delay, since the interconnection wire is not appropriately segmented by the repeaters. To solve these problems, we designed the IPs such that designers can embed the repeaters in the IP for the SoC integration. In other words, it allows the cross-IP interconnections to be routed over the IP using repeaters inserted in the IP. The design concept, physical implementation, and application examples of the embedded repeaters are described in this brief. Experimental results show that the proposed design will not only make the floor plan of the SoC easier but will also improve the signal delay and the power consumption of the long interconnection circuits. Jinn-Shyan Wang, Keng-Jui Chang, Chingwei Yeh, Shih-Chieh Chang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | Post silicon skew tuning: Survey and analysisabstractClock skew minimization is an important design consideration. However, with the advance of the technology and the smaller device scaling, Process, Voltage, and Temperature (PVT) variations make the clock skew minimization face great challenges. To mitigate the impact of PVT variations, many previous works proposed the Post Silicon Tuning (PST) architecture to dynamically balance the skew of a clock tree. In the PST architecture, there are two main components: Adjustable Delay Buffer (ADB) and Phase Detector (PD). In this paper, we make a survey about existing techniques to the PST architecture and introduce several important design concerns such as the ADB selection, system controlling, and design testing to the PST architecture. Mac Y. C. Kao, Kun-Ting Tsai, Hsuan-Ming Chou, Shih-Chieh Chang 0001 |
ASP-DAC | 4 |
| 2012 | A probabilistic analysis method for functional qualification under Mutation AnalysisabstractMutation Analysis (MA) is a fault-based simulation technique that is used to measure the quality of testbenches in error (mutant) detection. Although MA effectively reports the living mutants to designers, it suffers from the high simulation cost. This paper presents a probabilistic MA preprocessing technique, Error Propagation Analysis (EPA), to speed up the MA process. EPA can statically estimate the probability of the error propagation with respect to each mutant for guiding the observation-point insertion. The inserted observation-points will reveal a mutant's status earlier during the simulation such that some useless testcases can be discarded later. We use the mutant model from an industrial EDA tool, Certitude, to conduct our experiments on the OpenCores' RT-level designs. The experimental results show that the EPA approach can save about 14% CPU time while obtaining the same mutant status report as the traditional MA approach. Hsiu-Yi Lin, Chun-Yao Wang, Shih-Chieh Chang 0001, Yung-Chih Chen, Hsuan-Ming Chou, Ching-Yi Huang, Yen-Chi Yang, Chun-Chien Shen |
DATE | 3 |
| 2012 | Mitigating lifetime underestimation: A system-level approach considering temperature variations and correlations between failure mechanismsabstractLifetime (long-term) reliability has been a main design challenge as technology scaling continues. Time-dependent dielectric breakdown (TDDB), negative bias temperature instability (NBTI), and electromigration (EM) are some of the critical failure mechanisms affecting lifetime reliability. Due to the correlation between different failure mechanisms and their significant dependence on the operating temperature, existing models assuming constant failure rate and additive impact of failure mechanisms will underestimate the lifetime of a system, usually measured by mean-time-to-failure (MTTF). In this paper, we propose a new methodology which evaluates system lifetime in MTTF and relies on Monte-Carlo simulation for verifying results. Temperature variations and the correlation between failure mechanisms are considered so as to mitigate lifetime underestimation. The proposed methodology, when applied on an Alpha 21264 processor, provides less pessimistic lifetime evaluation than the existing models based on sum of failure rate. Our experimental results also indicate that, by considering the correlation of TDDB and NBTI, the lifetime of a system is likely not dominated by TDDB or NBTI, but by EM or other failure mechanisms. Kai-Chiang Wu, Ming-Chao Lee, Diana Marculescu, Shih-Chieh Chang 0001 |
DATE | 4 |
| 2012 | Efficient multiple-bit retention register assignment for power gated design: Concept and algorithmsabstractRetention registers have been widely used in power gated design to store data during sleep mode. Since they consume much larger area and power than normal registers, it is imperative to minimize the total retention storage size. The current industry practice only replace all registers with single-bit retention ones, which significantly limits the design freedom and results in excessive area and power overhead. Towards this, for the first time in literature, we propose the concept of multi-bit retention register, with which only selected registers need to be replaced. It can significantly reduce the number of bits that need to be stored and thus the area and leakage power, but needs several clock cycles for mode transition. In addition, an efficient assignment algorithm is developed to minimize the total retention storage size subject to mode transition latency constraint. Experimental results show that our framework on average can reduce the leakage power in sleep mode and the retention storage area by 66.03%, compared with the single-bit retention register based design. Yu-Guang Chen, Yiyu Shi 0001, Kuan-Yu Lai, Hui Geng, Shih-Chieh Chang 0001 |
ICCAD | 5 |
| 2012 | Memory-efficient pattern matching architectures using perfect hashing on graphic processing unitsabstractMemory architectures have been widely adopted in network intrusion detection system for inspecting malicious packets due to their flexibility and scalability. Memory architectures match input streams against thousands of attack patterns by traversing the corresponding state transition table stored in commodity memories. With the increasing number of attack patterns, reducing memory requirement has become critical for memory architectures. In this paper, we propose a novel memory architecture using perfect hashing to condense state transition tables without hash collisions. The proposed memory architecture achieves up to 99.5% improvement in memory reduction compared to the traditional two-dimensional memory architecture. We have implemented our memory architectures on graphic processing units and tested using attack patterns from Snort V2.8 and input packets form DEFCON. The experimental results show that the proposed memory architectures outperform state-of-the-art memory architectures both on performance and memory efficiency. Chen-Hsiung Liu, Shih-Chieh Chang 0001, Wing-Kai Hon |
INFOCOM | 3 |
| 2012 | Efficient on-line module-level wake-up scheduling for high performance multi-module designsabstractPower consumption has become the major bottleneck for modern high-performance architectures, which typically contain large numbers of modules. To suppress leakage power, sleep transistors have been extensively used, and wake-up scheduling is needed to determine the wake-up times and order of these sleep transistors. Most existing works on wake-up scheduling are based on sleep transistors and delay buffers in daisy-chains; they work well for the gate-level scheduling within a module when all the gates need to be turned on. Yet, for state-of-the-art designs, the number of modules that need to be turned on and their locations may vary depending on the task to be performed at runtime. Accordingly, we cannot extend the existing gate-level scheduling algorithms to decide the module-level wake-up order. To address the problem, we propose to first off-line construct a multi-conflict graph (MCG) based on the noise constraints; based on the graph, we then develop an on-line algorithm to decide the wake-up order. Experimental results show that on average, the wake-up latency from our approach is not only 46.01% shorter compared with the existing work but also conservatively only 0.45% longer than that from a Monte Carlo search-based evaluation, which is orders of magnitude slower. To the best of our knowledge, this is the first in-depth study on on-line module-level wake-up scheduling for high-performance architectures. Ming-Chao Lee, Yiyu Shi 0001, Yu-Guang Chen, Diana Marculescu, Shih-Chieh Chang 0001 |
ISPD | 5 |
| 2012 | Efficient Wakeup Scheduling Considering Both Resource Usage and Timing Budget for Power Gating DesignsabstractPower gating has been a very effective way to reduce power leakage. Normally, wakeup scheduling is required to control the turn-on times of sleep transistors to limit the surge current during the wakeup process. In this paper, a voltage sensor is adopted to compare the virtual ground voltage with the predesigned reference voltages, and use the result to determine the turn-on times of sleep transistors. We then propose a novel wakeup scheduling formulation that considers the tradeoff between wakeup times and hardware resources incurred by the voltage sensor. To address this problem, on the one hand, an efficient algorithm is proposed to find a wakeup scheduling with minimum wakeup time under the resource constraint. On the other hand, the algorithm to find a wakeup scheduling with minimum resource usage under the timing budget is presented. Experimental results show that with little increase in wakeup times, our algorithm can achieve significant hardware resource reduction for power gating designs. Ming-Chao Lee, Yiyu Shi 0001, Shih-Chieh Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2012 | Timing Optimization in Sequential Circuit by Exploiting Clock-Gating LogicabstractClock gating is a popular technique for reducing power dissipation. In a circuit with clock gating, the clock signal can be shut off without changing the functionality under certain clock-gating conditions. In this article, we observe that the clock-gating conditions and the next-state function of a Flip-Flop (FF) are correlated and can be used for sequential circuit optimization. We also show that the implementation of the next-state function of any FF can be just an inverter if the clock signal is appropriately gated. By exploiting the flexibility between the clock-gating conditions and the next-state function, we propose an iterative optimization algorithm to improve the timing of sequential circuits. We present experimental results of a set of benchmark circuits with a timing improvement of 10.20% on average. Shih-Hung Weng, Yu-Min Kuo, Shih-Chieh Chang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2011 | NBTI-aware power gating designabstractA header-based power gating structure inserts PMOS as sleep transistors between the power rail and the circuit. Since PMOS sleep transistors in the functional mode are turned-on continuously, Negative Bias Temperature Instability (NBTI) influences the lifetime reliability of PMOS sleep transistors seriously. To tolerate NBTI effect, sizes of PMOS sleep transistors are normally over-sized. In this paper, we propose a novel NBTI-aware power gating architecture to extend the lifetime of PMOS sleep transistors. In our structure, sleep transistors are switched on/off periodically so that overall turned-on times of sleep transistors are reduced and sleep transistors are less influenced by NBTI effect. The experimental results show that our approach can achieve better lifetime extensions of PMOS sleep transistors than previous works and few area overheads. Ming-Chao Lee, Yu-Guang Chen, Ding-Kei Huang, Shih-Chieh Chang 0001 |
ASP-DAC | 4 |
| 2011 | Fault-tolerant 3D clock networkabstractClock tree synthesis is one of the most important and challenging problems in 3D ICs. The clock signals have to be delivered by through-silicon vias (TSVs) to different tiers with minimum skew and latency. While there are a few related works in literature, none of them considers the reliability of TSVs. Accordingly, the failure of any TSV in the clock tree yields a bad chip. The naive solution using double-TSV can alleviate the problem. But the significant area overhead renders it less practical for large designs. In this paper, we propose a novel TSV fault-tolerant unit (TFU) that can provide tolerance against TSV failures in a 3D clock network. It makes use of the existing 2D redundant trees designed for pre-bond testing, and thus has minimum area overhead. Compared to the double TSV technique, the 3D clock network constructed by our TFUs can achieve 61% area reduction with 3.9% yield rate improvement on an industrial case. To the best of the authors' knowledge, this is the first practical work in literature that considers the fault tolerance of a 3D clock network. Chiao-Ling Lung, Yu-Shih Su, Shih-Hsiu Huang, Yiyu Shi 0001, Shih-Chieh Chang 0001 |
DAC | 5 |
| 2011 | Thermal-aware on-line task allocation for 3D multi-core processor throughput optimizationabstractThree-dimensional integrated circuit (3D IC) has become an emerging technology in view of its advantages in packing density and flexibility in heterogeneous integration. The multi-core processor (MCP), which is able to deliver equivalent performance with less power consumption, is a candidate for 3D implementation. However, when maximizing the throughput of 3D MCP, due to the inherent heat removal limitation, thermal issues must be taken into consideration. Furthermore, since the temperature of a core strongly depends on its location in the 3D MCP, a proper task allocation helps to alleviate any potential thermal problem and improve the throughput. In this paper, we present a thermal-aware on-line task allocation algorithm for 3D MCPs. The results of our experiments show that our proposed method achieves 16.32X runtime speedup, and 23.18% throughput improvement. These are comparable to the exhaustive solutions obtained from optimization modeling software LINGO. On average, our throughput is only 0.85% worse than that of the exhaustive method. In 128 task-to-core allocations, our method takes only 0.932 ms, which is 57.74 times faster than the previous work. Chiao-Ling Lung, Yi-Lun Ho, Ding-Ming Kwai, Shih-Chieh Chang 0001 |
DATE | 4 |
| 2011 | Accelerating Regular Expression Matching Using Hierarchical Parallel Machines on GPUabstractDue to the conciseness and flexibility, regular expressions have been widely adopted in Network Intrusion Detection Systems to represent network attack patterns. However, the expressive power of regular expressions accompanies the intensive computation and memory consumption which leads to severe performance degradation. Recently, graphics processing units have been adopted to accelerate exact string pattern due to their cost-effective and enormous power for massive data parallel computing. Nevertheless, so far as the authors are aware, no previous work can deal with several complex regular expressions which have been commonly used in current NIDSs and been proven to have the problem of state explosion. In order to accelerate regular expression matching and resolve the problem of state explosion, we propose a GPU-based approach which applies hierarchical parallel machines to fast recognize suspicious packets which have regular expression patterns. The experimental results show that the proposed machine achieves up to 117 Gbps and 81 Gbps in processing simple and complex regular expressions, respectively. The experimental results demonstrate that the proposed parallel approach not only resolves the problem of state explosion, but also achieves much more acceleration on both simple and complex regular expressions than other GPU approaches. Chen-Hsiung Liu, Shih-Chieh Chang 0001 |
GLOBECOM | 3 |
| 2011 | On the preconditioner of conjugate gradient method - A power grid simulation perspectiveabstractPreconditioned Conjugate Gradient (PCG) method has been demonstrated to be effective in solving large-scale linear systems for sparse and symmetric positive definite matrices. One critical problem in PCG is to design a good preconditioner, which can significantly reduce the runtime while keeping memory usage efficient. Universal preconditioners are simple and easy to construct, but their effectiveness is highly problem-dependent. On the other hand, domain-specific preconditioners that explore the underlying physical meaning of the matrices usually work better, but are difficult to design. In this paper, we study the problem in the context of power grid simulation, and develop a novel preconditioner based on the power grid structure through simple circuit simulations. Experimental results show 43% reduction in the number of iterations and 23% speedup over existing universal preconditioners. Chung-Han Chou, Nien-Yu Tsai, Hao Yu 0001, Che-Rung Lee, Yiyu Shi 0001, Shih-Chieh Chang 0001 |
ICCAD | 6 |
| 2011 | Useful-skew clock optimization for multi-power mode designsabstractInstead of minimizing clock skew, skew can be useful to improve circuit performance. However, it is difficult to apply useful skew to a design with complicated power modes. With only one clock tree, useful skew in one power mode may be harmful in another power mode. In this paper, we propose to use adjustable delay buffers (ADBs) to construct a tunable clock tree so that useful skew can be assigned for different power modes. Assuming positions of ADBs are determined, we assign delays of ADBs for each power mode by LP. Then a speedup theorem is proposed to greatly reduce LP inequalities. We also propose an efficient method to select positions of ADBs. Our experimental results show that average 99.45% inequities are decreased and an average performance improvement of 27.35% is obtained compared with commercial tool SOC Encounter™. Hsuan-Ming Chou, Shih-Chieh Chang 0001 |
ICCAD | 3 |
| 2011 | A robust architecture for post-silicon skew tuningabstractClock skew minimization is important in VLSI design field. Due to the presence of Process, Voltage, and Temperature (PVT) variations, the Post-Silicon Skew Tuning (PST) technique with the ability of tolerating PVT variations has brought a broad discussion. A PST architecture can dynamically minimize the clock skew even after a chip is manufactured. However, testing the variation tolerance ability of a PST architecture is very difficult because the clock skew does not directly affect the functionality of a design. In addition, creating PVT variation in the traditional testing environment is not easy. Unlike most previous works which focus on the implementation and the performance issues of a PST architecture, the objective of this paper is to propose efficient test mechanisms and verify the variation tolerance ability. In addition, we also propose a novel structure to increase the robustness of a PST architecture in case of a manufacturing fault. Our experiment shows that with little overhead, we can achieve robustness. Mac Y. C. Kao, Kun-Ting Tsai, Shih-Chieh Chang 0001 |
ICCAD | 3 |
| 2011 | Analysis and mitigation of NBTI-induced performance degradation for power-gated circuits
Kai-Chiang Wu, Diana Marculescu, Ming-Chao Lee, Shih-Chieh Chang 0001 |
ISLPED | 4 |
| 2011 | Efficient Pattern Matching Algorithm for Memory ArchitectureabstractNetwork intrusion detection system is used to inspect packet contents against thousands of predefined malicious or suspicious patterns. Because traditional software alone pattern matching approaches can no longer meet the high throughput of today's networking, many hardware approaches are proposed to accelerate pattern matching. Among hardware approaches, memory-based architecture has attracted a lot of attention because of its easy reconfigurability and scalability. In order to accommodate the increasing number of attack patterns and meet the throughput requirement of networks, a successful network intrusion detection system must have a memory-efficient pattern-matching algorithm and hardware design. In this paper, we propose a memory-efficient pattern-matching algorithm which can significantly reduce the memory requirement. For Snort rule sets, the new algorithm achieves 21% of memory reduction compared with the traditional Aho-Corasick algorithm. In addition, we can gain 24% of memory reduction by integrating our approach to the bit-split algorithm which is the state-of-the-art memory-based approach. Shih-Chieh Chang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | Performance Optimization Using Variable-Latency Design StyleabstractIn many designs, the worst-case delay of a critical path may be activated infrequently. Traditional optimization approaches assume the worst-case conditions, which could lead to an inefficient resource usage. It is possible to improve the throughput of such designs by introducing variable latency. One existing realization of the variable-latency design style is based on telescopic units. The design of the hold logic in telescopic units influences the circuit's throughput. In this paper, we show that the traditionally designed hold logic may be inaccurate. We use the short path activation conditions to obtain more accurate hold logic and improve the efficiency of telescopic units. To reduce the overhead for large circuits, we propose an efficient heuristic methodology of constructing non-exact hold logic. We also discuss how to choose the telescopic unit's timing constraint. On average, our approach achieves the performance gain of 21.67% compared to 13.99%, reported in the previous work. Yu-Shih Su, Da-Chung Wang, Shih-Chieh Chang 0001, Malgorzata Marek-Sadowska |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | An efficient phase detector connection structure for the skew synchronization systemabstractClock skew optimization continues to be an important concern in circuit designs. To overcome the influence caused by PVT variations, the automatic skew synchronization scheme can dynamically adjust and reduce the clock skew after a chip is manufactured. There are two key components in the skew synchronization scheme: Adjustable Delay Buffer (ADB) and Phase Detector (PD). Most previous researchers have emphasized on ADB placement issues. In this paper, we show that the connection between FFs and PDs can also greatly influence the final clock skew due to the insertion of the PDs. We first analyze the influence of PD connection structures. Then we propose an algorithm to generate a PD connection structure which achieves the minimum influence to the clock skew. Our experimental results are very encouraging. Yu-Chien Kao, Hsuan-Ming Chou, Kun-Ting Tsai, Shih-Chieh Chang 0001 |
DAC | 4 |
| 2010 | Clock skew optimization considering complicated power modesabstractTo conserve energy, a design which utilizes different power modes has been widely adopted. However, when a design has many different power modes, clock tree optimization (CTO) becomes very difficult. In this paper, we propose a two-level power-mode-aware CTO methodology. Among all different power modes, the chip-level CTO globally reduces clock skew among modules, whereas the module-level CTO reduces clock skew within a single module. Our experimental results show that the power-mode-aware CTO can achieve significant improvement in the worst-case condition with only a minor penalty in area. Chiao-Ling Lung, Zi-Yi Zeng, Chung-Han Chou, Shih-Chieh Chang 0001 |
DATE | 4 |
| 2010 | Accelerating String Matching Using Multi-Threaded Algorithm on GPUabstractNetwork Intrusion Detection System has been widely used to protect computer systems from network attacks. Due to the ever-increasing number of attacks and network complexity, traditional software approaches on uni-processors have become inadequate for the current high-speed network. In this paper, we propose a novel parallel algorithm to speedup string matching performed on GPUs. We also innovate new state machine for string matching, the state machine of which is more suitable to be performed on GPU. We have also described several speedup techniques considering special architecture properties of GPU. The experimental results demonstrate the new algorithm on GPUs achieves up to 4,000 times speedup compared to the AC algorithm on CPU. Compared to other GPU approaches, the new algorithm achieves 3 times faster with significant improvement on memory efficiency. Furthermore, because the new Algorithm reduces the complexity of the Aho-Corasick algorithm, the new algorithm also improves on memory requirements. Sheng-Yu Tsai, Chen-Hsiung Liu, Shih-Chieh Chang 0001, Jyuo-Min Shyu |
GLOBECOM | 4 |
| 2010 | Synthesis of an efficient controlling structure for post-silicon clock skew minimizationabstractClock skew minimization has been an important design constraint. However, due to the complexity of Process, Voltage, and Temperature (PVT) variations, the minimization of clock skew has faced a great challenge. To overcome the influence of PVT variations, several previous works proposed Post Silicon Tuning (PST) architecture to dynamically balance the skew of a clock tree. In the PST architecture, there are two main components: Adjustable Delay Buffer (ADB) and Phase Detector (PD). Most previous works focus on determining good positions of ADBs in a PST design. In this paper, we first show that which pairs of FFs are connected to PDs, called PD structure, also greatly influence the complexity of hardware control for a PST design. Without careful planning of a PD structure, we need large number of control signals to adjust the delays of ADBs. In addition, we also show that a PD structure may influence the accuracy of the clock skew. Among possible connection structures, this paper proposes an efficient PD structure which not only simplifies the hardware control but also minimizes the clock skew of a PST design. Yu-Chien Kao, Hsuan-Ming Chou, Kun-Ting Tsai, Shih-Chieh Chang 0001 |
ICCAD | 4 |
| 2010 | Sleep Transistor Sizing for Leakage Power Minimization Considering Temporal CorrelationabstractPower gating is one of the most effective ways to reduce leakage power. In this paper, we introduce a new relationship among maximum instantaneous current, IR-drops and sleep transistor networks from a temporal viewpoint. Based on this relationship, we propose an algorithm to reduce the total sizes of sleep transistors in distributed sleep transistor network designs with the consideration of decoupling capacitances is taken. Our method achieves significantly better results than previous works on sleep transistor sizes. De-Shiuan Chiou, Yuting Chen 0002, Da-Cheng Juan, Shih-Chieh Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2010 | Clock Skew Minimization in Multi-Voltage Mode Designs Using Adjustable Delay BuffersabstractIn synchronous circuit designs, clock skew is difficult to minimize because a single physical layout of a clock tree must satisfy multiple constraints in a complicated power mode environment where certain modules may operate with different voltages. In this paper, we use adjustable delay buffers (ADB) whose delays can be tuned or adjusted to minimize clock skew under different power modes. Assuming that the positions of$k$ADBs are already determined, we first propose a linear-time optimal algorithm which assigns the values of ADBs so that the skew is optimal among all possible ADB assignments with a possibility of latency penalty. Then, we propose a modified optimal algorithm without latency penalty. We also propose an efficient heuristic to determine good positions for ADBs. Our results show significant improvement when compared to cases without ADBs. Yu-Shih Su, Wing-Kai Hon, Cheng-Chih Yang, Shih-Chieh Chang 0001, Yeong-Jar Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2010 | An Efficient Wake-Up Strategy Considering Spurious Glitches Phenomenon for Power Gating DesignsabstractDuring the power mode transition, simultaneously turning on sleep transistors provides a sufficiently large surge current, which may cause a large IR drop in the power networks. The IR drop in turn causes errors in the retention sequential elements of the sleep modules or errors of the nonsleep modules. One efficient way to control the surge current is to schedule the turn-on sequences of sleep transistors. In this paper, we introduce several important properties of the surge current during the power mode transition for the distributed sleep transistor network (DSTN) design, which is a popular power gating design style. Based on these properties, we propose an accurate estimation of the surge current and provide efficient schedules on the DSTN structure. Our methods achieved significantly better results than previous works-on average, 261 times wake-up time reduction and 30% less energy loss during the power mode transition. Da-Cheng Juan, Yuting Chen 0002, Ming-Chao Lee, Shih-Chieh Chang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2009 | An efficient wakeup scheduling considering resource constraint for sensor-based power gating designsabstractPower gating has been a very effective way to reduce leakage power. One important design issue for a power gating design is to limit the surge current during the wakeup process. Normally, a wakeup scheduling is required to control turn-on times of sleep transistors. In this paper, we adopt a voltage sensor to compare pre-designed reference voltages with the virtual ground voltage and use the comparison result to determine turn-on times of sleep transistors. Special properties and optimizations of using voltage sensors are discussed. Since a wakeup scheduling with fast wakeup time may require significant hardware resources, we propose a new wakeup scheduling formulation which considers the trade-off between wakeup times and hardware resources. Our experimental results show that with small increases on wakeup times, we can reduce significant hardware resources for a power gating design. Ming-Chao Lee, Yuting Chen 0002, Yo-Tzu Cheng, Shih-Chieh Chang 0001 |
ICCAD | 4 |
| 2009 | Value assignment of adjustable delay buffers for clock skew minimization in multi-voltage mode designsabstractIn synchronous circuit designs, clock skew is difficult to minimize because a single physical layout of a clock tree must satisfy multiple constraints in a complicated power mode environment where certain modules may operate with different voltages. In this paper, we use Adjustable Delay Buffers (ADB) whose delays can be tuned or adjusted to minimize clock skew under different power modes. Assuming that the positions of k ADBs are already determined, we propose a linear-time optimal algorithm which assigns the values of ADBs so that the skew is optimal among all possible ADB assignments. We also propose an efficient heuristic to determine good positions for ADBs. Our results show significant improvement when compared to cases without ADBs. Categories and Subject Descriptors B.6.3 [Logic Design]: Design Aids — Optimization General Terms: Algorithms, Reliability Yu-Shih Su, Wing-Kai Hon, Cheng-Chih Yang, Shih-Chieh Chang 0001, Yeong-Jar Chang |
ICCAD | 4 |
| 2009 | Efficient Boolean Characteristic Function for Timed Automatic Test Pattern GenerationabstractTiming analysis is critical for many circuit optimizations. An accurate timing analysis can be achieved by finding input vectors that simultaneously satisfy both functional and temporal requirements. The problem of finding such input vectors can be modeled as a Boolean equation called the timed characteristic function (TCF). Despite the usefulness of the TCF, traditional TCF construction and solving is slow for large circuits. In this paper, we present a more efficient way to use the TCF. On average, our method is much faster than other most recent works. Yu-Min Kuo, Yue-Lung Chang, Shih-Chieh Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2009 | Spare Cells With Constant Insertion for Engineering ChangeabstractEngineering change (EC) is the process of modifying a VLSI design implementation to eliminate design errors, to add new specifications, or to correct design constraint violations. Usually, an EC problem is resolved by using spare cells that have been inserted into unused spaces on a chip. In this paper, we describe an iterative method to determine feasible mapping solutions for an EC problem considering spare cells whose inputs can be connected toVddorGnd. Setting some of the cell inputs to fixed values is referred to asconstantinsertion. Constant insertion can increase cells' functional flexibility. Our experimental results suggest that constant insertion reduces the area required to find a feasible mapping solution to 80% of that with no constant insertion for the selected EC equations. We also show a procedure for modifying the initial feasible EC solution such that the routing or timing improves. Yu-Min Kuo, Ya-Ting Chang, Shih-Chieh Chang 0001, Malgorzata Marek-Sadowska |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2009 | Sleep Transistor Sizing for Leakage Power Minimization Considering Charge BalancingabstractOne of the effective techniques to reduce leakage power is power gating. Previously, a distributed sleep transistor network was proposed to reduce the sleep transistor area for power gating by connecting all the virtual ground lines together to minimize the maximum instantaneous current flowing through sleep transistors. In this paper, we propose a new methodology for determining the sizes of sleep transistors of the DSTN structure. We present novel algorithms and theorems for efficiently estimating a tight upper bound of the voltage drop and minimizing the sizes of sleep transistors. We also present mathematical proofs of our theorems and lemmas in detail. Our experimental results show 23.36% sleep transistor area reduction compared to the previous work on average. De-Shiuan Chiou, Shih-Hsin Chen, Shih-Chieh Chang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2008 | Behavioral synthesis with activating unused flip-flops for reducing glitch power in FPGAabstractIn this paper we discuss optimizing the interconnect power of designs implemented in FPGA platforms. In particular, we reduce the glitch power on interconnects associated with the output of functional units in a design. The idea is to activate unused flip-flops to block the propagation of glitches, which takes advantage of the abundant flip-flops in modern FPGA structures. Since the activation of additional flip-flops may cause data hazard problems, we develop several effective behavioral synthesis techniques to prevent such data hazards. We also study the optimality of our techniques. The experimental results show that on average, our methods lead to a 28% reduction in dynamic power in the Xilinx Virtex-II platform. Cheng-Tao Hsieh, Jason Cong, Zhiru Zhang, Shih-Chieh Chang 0001 |
ASP-DAC | 4 |
| 2008 | A novel sequential circuit optimization with clock gating logicabstractTo save power consumption, it has been shown that the clock signal can be gated without changing the functionality under certain clock-gating conditions. We observe that the clock-gating conditions and the next-state function of a Flip-Flop (FF) are correlated and can be used for sequential optimization. We show that the implementation of the next-state function of any FF can be just an inverter if the clock signal is appropriately gated. By exploiting the flexibility between the clock-gating conditions and the next-state function, we propose an iterative optimization technique to minimize the overall timing. Yu-Min Kuo, Shih-Hung Weng, Shih-Chieh Chang 0001 |
ICCAD | 3 |
| 2008 | Timing analysis considering IR drop waveforms in power gating designsabstractIR drop noise has become a critical issue in advanced process technologies. Traditionally, timing analysis in which the IR drop noise is considered assumes a worst-case IR drop for each gate; however, using this assumption provides unduly pessimistic results. In this paper, we describe a timing analysis approach for power gating designs. To improve the accuracy of the gate delay calculation we determine the virtual voltage level by taking into account the IR drop waveforms across the sleep transistors. These can be obtained efficiently using a linear programming approach. Our experimental results are very promising. Shih-Hung Weng, Yu-Min Kuo, Shih-Chieh Chang 0001, Malgorzata Marek-Sadowska |
ICCD | 3 |
| 2008 | Synthesis of a novel timing-error detection architectureabstractDelay variation can cause a design to fail its timing specification. Ernst et al. [2003] observe that the worst delay of a design is least probable to occur. They propose a mechanism to detect and correct occasional errors while the design can be optimized for the common cases. Their experimental results show significant performance (or power) gain as compared with the worst-case design. However, the architecture in Ernst et al. [2003] suffers the short path problem, which is difficult to resolve. In this article, we propose a novel error-detecting architecture to solve the short path problem. Our experimental results show considerable performance gain can be achieved with reasonable area overhead. Yu-Shih Su, Po-Hsien Chang, Shih-Chieh Chang 0001, TingTing Hwang |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2007 | Optimization of pattern matching algorithm for memory based architectureabstractDue to the advantages of easy re-configurability and scalability, the memory-based string matching architecture is widely adopted by network intrusion detection systems (NIDS). In order to accommodate the increasing number of attack patterns and meet the throughput requirement of networks, a successful NIDS system must have a memory-efficient pattern-matching algorithm and hardware design. In this paper, we propose a memory-efficient pattern-matching algorithm which can significantly reduce the memory requirement. For total Snort string patterns, the new algorithm achieves 29% of memory reduction compared with the traditional Aho-Corasick algorithm [5]. Moreover, since our approach is orthogonal to other memory reduction approaches, we can obtain substantial gain even after applying the existing state-of-the-art algorithms. For example, after applying the bit-split algorithm [9], we can still gain an additional 22% of memory reduction. Yu-Tang Tai, Shih-Chieh Chang 0001 |
ANCS | 3 |
| 2007 | Fine-Grained Sleep Transistor Sizing Algorithm for Leakage Power MinimizationabstractPower gating is one of the most effective ways to reduce leakage power. In this paper, we introduce a new relationship among Maximum Instantaneous Current, IR drops and sleep transistor networks from a temporal viewpoint. Based on this relationship, we propose an algorithm to reduce the total sizes of sleep transistors in Distributed Sleep Transistor Network designs. On average, the proposed method can achieve 21% reduction in the sleep transistor size. De-Shiuan Chiou, Da-Cheng Juan, Yuting Chen 0002, Shih-Chieh Chang 0001 |
DAC | 4 |
| 2007 | An Efficient Mechanism for Performance Optimization of Variable-Latency DesignsabstractIn many designs, the worst-case-delay path may never be exercised or may be exercised infrequently. For those designs, a strategy of optimizing a circuit for the worst-case conditions could lead to inefficient resource use. It is possible to improve the throughput of such circuits by introducing variable latency. One of the existing realizations of variable-latency design style is based on Telescopic Units. The design of the hold logic in telescopic units influences the circuit's throughput. In this paper, we show that the traditionally-designed hold logic in telescopic units may be inaccurate. We make use of the short path activation conditions to obtain more accurate hold logic than that commonly applied in the telescopic units. On average, our approach achieves a performance gain of 25.79% compared to 14.04%, which was reported in the previous works. Yu-Shih Su, Da-Chung Wang, Shih-Chieh Chang 0001, Malgorzata Marek-Sadowska |
DAC | 3 |
| 2007 | An efficient wake-up schedule during power mode transition considering spurious glitches phenomenonabstractDuring the power mode transition, a large surge current may lead to the malfunctions in a power-gating design. In this paper, we introduce several important properties of the surge current during the power mode transition for the distributed sleep transistor network (DSTN) designs. Based on these properties, we propose an accurate estimation of surge current and provide an efficient schedule on the DSTN structure. Our experiment achieved significantly better results than previous works - on average, 332 times wake-up time reduction and 35.48% less energy loss during the power mode transition. Yuting Chen 0002, Da-Cheng Juan, Ming-Chao Lee, Shih-Chieh Chang 0001 |
ICCAD | 4 |
| 2007 | Engineering change using spare cells with constant insertionabstractIn the VLSI design process, a design implementation often needs to be corrected because of new specifications or design constraint violations. This correction process is referred to as engineering change (EC). Usually, an EC problem is resolved by using spare cells, which have been inserted into the unused spaces of a chip. In this paper, we propose an iterative method to generate feasible mapping solutions for an EC problem considering spare cells whose inputs may be tied to Vdd or Gnd, called constant insertion. Applying constant insertion can increase a cell's flexibility in aspect of functionalities, so far-away spare cells need not be used just for some specific functionality. Our experimental results show that the area in which there are enough spare cells for a mapping solution with constant insertion is only 82% of the area without constant insertion. Yu-Min Kuo, Ya-Ting Chang, Shih-Chieh Chang 0001, Malgorzata Marek-Sadowska |
ICCAD | 3 |
| 2007 | Analysis and optimization of power-gated ICs with multiple power gating configurationsabstractPower gating is an efficient technique for reducing leakage power in electronic devices by disconnecting blocks idle for long periods of time from the power supply. Disconnecting gated blocks causes changes in densities of currents flowing through a grid. Even in DC conditions, current densities in some grid branches may increase for some gating configurations to the extent of violating electromigration (EM) constraints. The existing DC methods for grid sizing optimize the grid area under voltage drop (IR) and EM constraints for one configuration of circuit blocks connected to the grid. We show that these methods cannot be directly applied for optimizing power-gated grids. We analyze the effects of EM and IR voltage drop in power grids with multiple power gating configurations. Based on our analyses, we develop a grid sizing algorithm to satisfy all reliability constraints for all feasible gating configurations. Our experimental results indicate that a grid initially sized for all blocks present may be modified to fulfill EM and IR constraints for multiple gating schedules with only a small area increase. Aida Todri, Malgorzata Marek-Sadowska, Shih-Chieh Chang 0001 |
ICCAD | 3 |
| 2007 | Electromigration and voltage drop aware power grid optimization for power gated ICsabstractPower gating is an efficient technique for reducing leakage power by disconnecting idle blocks from power supply. Gated blocks cause changes in current densities on the grid. Even in DC conditions for some power gating configuration (PGC), current densities in some branches may increase to the extent of violating electromigration (EM) constraints. The existing DC methods optimize the grid under voltage drop (IR) and EM constraints for a single configuration of blocks. We analyze the effects of power gating and develop a grid sizing algorithm to satisfy all reliability constraints for multiple PGCs with only a small increase in area. Aida Todri, Shih-Chieh Chang 0001, Malgorzata Marek-Sadowska |
ISLPED | 2 |
| 2007 | Optimization of Pattern Matching Circuits for Regular Expression on FPGAabstractRegular expressions are widely used in the network intrusion detection system (NIDS) to represent attack patterns. Previously, many hardware architectures have been proposed to accelerate regular expression matching using field-programmable gate array (FPGA) because FPGAs allow updating of new attack patterns. Because of the increasing number of attacks, we need to accommodate a large number of regular expressions on FPGAs. Although the minimization of logic equations has been studied intensively in the area of computer-aided design (CAD), the minimization of multiple regular expressions has been largely neglected. This paper presents a novel sharing architecture allowing our algorithm to extract and share common subregular expressions. Experimental results show that our sharing scheme significantly reduces the area of pattern matching circuits for regular expression. Chih-Tsun Huang, Chang-Ping Jiang, Shih-Chieh Chang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2006 | Delay variation tolerance for domino circuitsabstractFactors of delay variation, such as process variation and noise effects, may cause a manufactured chip to violate the pre-specified timing constraint. In this paper, we propose a re-synthesis technique to tolerate delay variation for domino circuits. Note that the slacks of nodes along critical paths are zero; any delay addition to those zero-slack nodes worsens the final performance of a circuit. Our basic idea is to increase the slacks of nodes in the critical region by appending a redundant auxiliary subcircuit to the original circuit. The auxiliary subcircuit can cause critical paths to become false paths or imperceptible paths as stated in S. Raj et al. (2004) so as to improve the capability of delay variation tolerance. Experimental results are very encouraging. Kai-Chiang Wu, Cheng-Tao Hsieh, Shih-Chieh Chang 0001 |
ASP-DAC | 3 |
| 2006 | Timing driven power gatingabstractPower Gating is effective for reducing leakage power. Previously, a Distributed Sleep Transistor Network (DSTN) was proposed to reduce the sleep transistor area by connecting all the virtual ground lines together to minimize the Maximum Instantaneous Current (MIC) through sleep transistors. In this paper, we propose a new methodology for determining the size of sleep transistors for the DSTN structure. We present novel algorithms and theorems for efficiently estimating a tight upper bound of the voltage drop. We also present efficient heurists for minimizing the sizes of sleep transistors. Our experimental results are very exciting. De-Shiuan Chiou, Shih-Hsin Chen, Shih-Chieh Chang 0001, Chingwei Yeh |
DAC | 3 |
| 2006 | Efficient Boolean characteristic function for fast timed ATPGabstractCircuit timing analysis is important in various aspects of circuit optimization. The problem of finding input vectors achieving functional and temporal requirements is known as timed Automatic Test Pattern Generation (timed ATPG). A timed ATPG algorithm will return an input vector that satisfies functional and temporal requirements simultaneously when evaluated. Several previous works use timed ATPG as a core engine for solving problems related to timing analysis, such as crosstalk and maximum instantaneous current analysis. Despite the usefulness of timed ATPG, traditional timed ATPG is slow and unscalable for large circuits. In this paper, we present a very efficient way for timed ATPG. On average, our results are 8 times faster than the most recent work, and in some cases, up to 32 times faster. Yu-Min Kuo, Yue-Lung Chang, Shih-Chieh Chang 0001 |
ICCAD | 3 |
| 2006 | Vectorless Estimation of Maximum Instantaneous Current for Sequential CircuitsabstractLarge current in a chip can cause problems such as noise and power consumption. In this paper, a vectorless approach to analyzing a tight upper bound on the maximum instantaneous current (MIC) of a circuit is proposed. Several types of signal correlations that can cause the MIC estimation to lose accuracy are first described. Next, taking signal correlations into account, theorems to identify gates that switch mutually exclusively are proposed. In particular, the proposed algorithm can naturally consider signal correlations across sequential elements (flip-flops), whereas previous research on this topic addressed combinational circuits only. After deriving the information of mutually exclusive switching, a graph algorithm is applied to obtain an upper bound on the MIC. On average, the obtained sequential benchmark results are 179% tighter than those from the iMax algorithm and 66% tighter than those from the partial input enumeration algorithm C.-T. Hsieh, J.-C. Lin, Shih-Chieh Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2006 | Multiple wire reconnections based on implication flow graphabstractGlobal flow optimization (GFO) can perform multiple fanout/fanin wire reconnections at a time by modeling the problem of multiple wire reconnections with a flow graph, and then solving the problem using the maxflow-mincut algorithm on the flow graph. In this article, we propose an efficient multiple wire reconnection technique that modifies the framework of GFO, and as a result, can obtain better optimization quality. First, we observe that the flow graph in GFO cannot fully characterize wire reconnections, which causes the GFO to lose optimality in several obvious cases. In addition, we find that fanin reconnection can have more optimization power than fanout reconnection, but requires more sophisticated modeling. We reformulate the problem of fanout/fanin reconnections by a new graph, called the implication flow graph (IFG). We show that the problem of wire reconnections on the implication flow graph is NP-complete and also propose an efficient heuristic on the new graph. To demonstrate the effectiveness of our proposed method, we conduct an application which utilizes the flexibility of the wire reconnections explored in the logic domain to further minimize interconnects in the physical layout. Our experimental results are very exciting. Zhong-Zhen Wu, Shih-Chieh Chang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2006 | Power minimization for dynamic PLAsabstractDynamic programmable logic arrays (PLAs) which are built of the nor-nor structure, have been very popular in high performance design because of their high-speed and predictable routing delay. However, the nor-nor structure incurs high switching activity in product lines and, thus, results in large power consumption. In this paper, we propose a new dynamic PLA structure which incorporates super product lines. A super product line adds the nand functionality on top of the nor structure, thus, lowering the switching activities in the product lines, as well as power consumption. Since there are many candidates for super product lines, we have developed a computer-aided design (CAD) algorithm based on the maximum weighted matching to find the optimal solution. We have performed experiments on a large set of Microelectronics Center of North Carolina (MCNC) benchmark circuits. The post simulation results show significant reduction in power consumption. Among the experimental circuits, circuit alu3 has the highest power saving 62.9% with the delay overhead 5.4%, and circuit newpla2 has the lowest power saving with delay overhead 22.7%. In addition, circuit in4 improves the delay with 5.7%. On the average, the power consumption can be saved 55.8% and the delay overhead is merely 3.3% for 25 circuits. Tzyy-Kuen Tien, Chih-Shen Tsai, Shih-Chieh Chang 0001, Chingwei Yeh |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2005 | Power estimation starategies for a low-power security processorabstractIn this paper, we present the power estimation methodologies for the development of a low-power security processor that contains significant amount of logic and memory. For the logic part, we present a highly accurate tool, called PowerMixer. This tool is a refinement of the so-called mixed-level methodology that combines the accuracy of quick SPICE and the speed of gate-level simulation. A grouping scheme is proposed so as to improve the accuracy for design blocks as large as 100K gates. For the memory part, we investigated the power consuming behavior of memories and point out the potential problems associated with the current commercial design flow. These tools, along with a previously published static peak power estimation method [4], jointly provide an evaluation platform for the power optimization and verification process of our security processor in a practical way. Yen-Fong Lee, Shi-Yu Huang, Sheng-Yu Hsu, I-Ling Chen, Cheng-Tao Shieh, Jian-Cheng Lin, Shih-Chieh Chang 0001 |
ASP-DAC | 7 |
| 2005 | Design and design automation of rectification logic for engineering changeabstractIn a later stage of a VLSI design, it is quite often to modify a design implementation to accommodate the new specification, design errors, or to meet design constraints. In addition to meet the design schedule for the new implementation, the reduction of the mask set have become very critical. In this paper, we propose a new method to add a programmable rectification module to reduce the mask cost and to improve the turn around time. When a modification is needed, one can program the rectification module to achieve the new implementation. The rectification module can be designed by one mask programmable gate array, or an embedded FPGA. To reduce the size needed for the rectification module, we also propose algorithms, which can intelligently select some internal signals of the old implementation to become pseudo primary inputs and primary outputs. Our experimental results are very encouraging. Yung-Chang Huang, Shih-Chieh Chang 0001, Wen-Ben Jone |
ASP-DAC | 3 |
| 2005 | Power minimization for dynamic PLAsabstractIn this paper, we propose a new dynamic PLA structure which incorporates super product lines. A super product line adds the NAND functionality on top of the NOR structure, thus lowering the switching activities in the product lines as well as power consumption. Since there are many candidates for super product lines, we have developed a CAD algorithm based on the maximum weighted matching to find optimal solution. The post simulation results show significant reduction in power consumption. On the average, the power consumption can be saved 58.9% and the delay overhead is merely 1.6% for 18 circuits. Tzyy-Kuen Tien, Chih-Shen Tsai, Shih-Chieh Chang 0001, Chingwei Yeh |
ASP-DAC | 3 |
| 2004 | Scan Chain Fault Identification Using Weight-Based Codes for SoC CircuitsabstractRecently, it has been observed that embedded cores in a high-speed SoC circuit have the problem of broken scan chains that cannot shift properly. Also, scan chain intermittent faults caused by hold-time violations and crosstalk noises are pervasive. In this research, an efficient method is proposed to identify the faulty scan chain(s) at the core level. That is, the core where the scan chain is defective can be identified, even if the scan chain is broken. The result can be used to tune up the fabrication process or to guide the fine-grained scan cell identification process. Here, weight-based m-out-of-n codes, which can generate a large number of codewords, with small hardware overhead and high fault detection capability are used to generate the scan chain diagnostic patterns for permanent (and possibly intermittent) faults. An efficient codeword generation method is proposed to maximize the number of codewords, minimize the aliasing probabilities and test application cost. The idea of multiple m-out-of-n codes is also proposed to guarantee that sufficient number of codewords are generated to perturb the scan chains and the associated combinational circuits. Simulation results demonstrate the feasibility of the proposed method. Swaroop Ghosh, K. W. Lai, Wen-Ben Jone, Shih-Chieh Chang 0001 |
Asian Test Symposium | 4 |
| 2004 | Re-synthesis for delay variation toleranceabstractSeveral factors such as process variation, noises, and delay defects can degrade the reliabilities of a circuit. Traditional methods add a pessimistic timing margin to resolve delay variation problems. In this paper, instead of sacrificing the performance, we propose a re-synthesis technique which adds redundant logics to protect the performance. Because nodes in the critical paths have zero slacks and are vulnerable to delay variation, we formulate the problem of tolerating delay variation to be the problem of increasing the slacks of nodes. Our re-synthesis technique can increase the slacks of all nodes or wires to be larger than a pre-determined value. Our experimental results show that additional area penalty is around 21% for 10% of delay variation tolerance. Shih-Chieh Chang 0001, Cheng-Tao Hsieh, Kai-Chiang Wu |
DAC | 1 |
| 2004 | A vectorless estimation of maximum instantaneous current for sequential circuitsabstractBoth the IR drop and EM problems require accurate analysis of maximum instantaneous current (MIC) on the power supply bus. We propose a vectorless approach to deriving a tight upper bound on MIC. We first characterize different types of signal correlation which may cause the MIC estimation to lose accuracy. We next propose theorems to identify gates which switch mutually exclusively, taking into account correlation across sequential elements (flip-flops). Note that previous research of this topic addressed on combinational circuits only. After obtaining the information of mutually exclusive switching, we then apply a graph algorithm to obtain an upper bound on MIC. In average, our results on sequential benchmarks are 212% tighter than those from iMax and 129% tighter than those from PIE based on H. Kriplani et al. (1995). Cheng-Tao Hsieh, Jian-Cheng Lin, Shih-Chieh Chang 0001 |
ICCAD | 3 |
| 2003 | Embedded core test generation using broadcast test architecture and netlist scramblingabstractIn this work, based on the concept of test pattern broadcasting, we propose a new core-based testing method which gives core users the maximum level of test freedom. Instead of only using the test patterns delivered by core providers, core users are allowed to broadcast their own test patterns to the cores of a SoC (system on chip) design for parallel scan testing. The fault coverage of each core test, using test patterns developed by any core user, can be evaluated by an enhanced version of a traditional fault simulator. The netlist of each core is scrambled before it is delivered to core users, thus the netlist will not be revealed. The enhanced fault simulator of a core has the capabilities of decoding the scrambled netlist, and performing fault simulation for the test patterns provided by each of the core users. For each core, both random test patterns (applied by a core user), and golden test patterns (delivered by the core provider) jointly achieve high and flexible fault coverage requirements. The enhanced logic simulator of each core can also decrypt the scrambled netlist, and perform logic simulation with the objective of generating fault-free test responses for signature analysis (for example). The proposed method has the advantages of minimizing the number of scan pins, reducing the test application time, and achieving the maximum level of test quality control by core users. Simulation results demonstrate the feasibility of this method. J. H. Jiang, Wen-Ben Jone, Shih-Chieh Chang 0001, Swaroop Ghosh |
IEEE Trans. Reliab. | 3 |
| 2002 | Crosstalk Alleviation for Dynamic PLAsabstractThe dynamic PLA style has become popular in designing high performance microprocessors because of its high speed and predictable routing delay. However like all other dynamic circuits, dynamic PLAs have suffered from the crosstalk noise problem. In this paper, we propose two techniques to alleviate crosstalk noise for dynamic PLAs. The first technique makes use of the fact that depending on the ordering of product lines, some crosstalk does not cause errors in outputs. A proper ordering can greatly reduce the number of lines affected by crosstalk noise. For those product lines which can be affected by crosstalk, we attempt to reduce the parallel length by re-ordering the input and output lines. We have performed experiments on a large set of MCNC benchmark circuits. The results show that after re-ordering, 86.7% of product lines become crosstalk immune and need not be considered for crosstalk prevention. Tzyy-Kuen Tien, Tong-Kai Tsai, Shih-Chieh Chang 0001 |
DATE | 3 |
| 2002 | Crosstalk alleviation for dynamic PLAsabstractThe dynamic programmable logic array (PLA) style has become popular in designing high-performance microprocessors because of its high speed and predictable routing delay. However, like all other dynamic circuits, dynamic PLAs have suffered from the crosstalk noise problem. The main reason is that the regularity of the PLA design style causes a product line parallel to the adjacent product lines on the same layer for a long distance so that the crosstalk noise can be significant. Most previous works attempt to prevent the crosstalk noise by adding additional devices or spacing between wires. However, the prevention may degrade performance and increase area/power. In this paper, the authors propose two techniques to alleviate crosstalk noise for dynamic PLAs. The first technique makes use of the fact that depending on the ordering of product lines, some crosstalk does not cause errors in outputs. A proper ordering can greatly reduce the number of product lines affected by crosstalk noise. The authors also observe that the parallel length between two adjacent lines depends on the input and output ordering. For those product lines which can be affected by crosstalk, they attempt to reduce the parallel length by reordering the input and output lines. They have performed experiments on a large set of MCNC benchmark circuits. The results show that after reordering, 86% of product lines become crosstalk immune and need not be considered for crosstalk prevention. Tzyy-Kuen Tien, Shih-Chieh Chang 0001, Tong-Kai Tsai |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2001 | Charge-sharing alleviation and detection for CMOS domino circuitsabstractCharge sharing, which occurs in any complementary metal-oxide-semiconductor (CMOS) domino gate, may degrade the output voltage level or may even cause an erroneous output value. In this paper, this problem is thoroughly investigated by considering circuit topology and circuit function. We describe a method to measure the sensitivity [called charge-sharing (CS) vulnerability] of the CS problem for each domino gate. A method to derive the CS vulnerability and the test vector for each domino gate is suggested. We also propose a transistor reordering method to dramatically reduce the CS vulnerabilities for all domino gates so that the CS problem can be alleviated. We also prove theoretically that a set of test vectors generated for single charge-sharing faults (SCSFs) can also detect all multiple charge-sharing faults (MCSFs). This good property significantly guarantees the test quality for the CS faults of domino circuits. Shih-Chieh Chang 0001, Ching-Hwa Cheng, Wen-Ben Jone, Shin-De Lee, Jinn-Shyan Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2001 | A timing-driven pseudoexhaustive testing for VLSI circuitsabstractBecause of its ability to detect all nonredundant combinational faults, exhaustive testing, which applies all possible input combinations to a circuit, is an attractive test method. However, the test application time for exhaustive testing can be very large. To reduce the test time, pseudoexhaustive testing inserts some bypass storage cells (bscs) so that the dependency of each node is within some predetermined value. Though bsc insertion can reduce the test time, it may increase circuit delay, In this paper, our objective is to reduce the delay penalty of bsc insertion for pseudoexhaustive testing. We first propose a tight delay lower bound algorithm, which estimates the minimum circuit delay for each node after bsc insertion. By understanding how the lower bound algorithm loses optimality, ne can propose a bsc insertion heuristic that tries to insert bscs so that the final delay is as close to the lower bound as possible. Our experiments show that the results of our heuristic are either optimal because they are the same as the delay lower bounds or they are very close to the optimal solutions. Shih-Chieh Chang 0001, Jiann-Chyi Rau |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2001 | Theorems and extensions of single wire replacementabstractIn this paper, we discuss the theorems and extensions of single alternative wire that attempts to replace one wire by another wire without changing the logic functionality. The wire replacement technique has been successfully applied to achieve logic optimization and routability improvement. However, there still exist several fundamental problems that have not been addressed such as whether the algorithm can find all single alternative wires. First, we present some cases of alternative wires, which the previous work (Chang et al., 1997) cannot obtain. Then, several theorems of tight necessary conditions and dominating conditions for a wire to be an alternative wire are proposed. With these theorems, we are able to derive an efficient procedure to find all possible alternative wires. The experimental results are very encouraging. Shih-Chieh Chang 0001, Zhong-Zhen Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2000 | Charge sharing fault analysis and testing for CMOS domino logic circuitsabstractBecause domino logic design offers smaller area and faster delay than conventional CMOS design, it is very popular in the high-performance processor. However, domino logic suffers from several problems and one of the most notable ones is the charge sharing problem. In this paper, we describe a method to measure the sensitivity of the charge-sharing problem for each domino gate. In addition, our algorithm also generates test vectors to detect the worst case of charge-sharing fault. Ching-Hwa Cheng, Wen-Ben Jone, Jinn-Shyan Wang, Shih-Chieh Chang 0001 |
Asian Test Symposium | 4 |
| 2000 | Novel techniques for improving testability analysisabstractThe purpose of a testability analysis program is to estimate the difficulty of testing a fault. A good measurement can give an early warning about the testing problem so as to provide guidance in improving the testability of a circuit. There have been researches attempting to efficiently compute the testability analysis. Among those, the Controllability and Observability Procedure COP (1984) can calculate the testability value of a stuck-at fault efficiently in a tree-structured circuit but may be very inaccurate for a general circuit. The inaccuracy in COP is due to the ignorance of signal correlations. The algorithm of TAIR (Testability Analysis by Implication Reasoning) proposes a testability analysis algorithm, which starts from the result of COP and then gradually improves the result by applying a set of rules. The set of rules in TAIR can capture some signal correlations and therefore the results of TAIR are more accurate than COP. In this paper, we first prove that the rules in TAIR can be replaced by a closed-form formulation. Then, based on the closed-form formulation, we proposed two novel techniques to further improve the testability analysis results. Our experimental results have shown improvement over the results of TAIR. Yin-He Su, Ching-Hwa Cheng, Shih-Chieh Chang 0001 |
Asian Test Symposium | 3 |
| 2000 | Wire Reconnections Based on Implication Flow GraphabstractGlobal Flow Optimization (GFO) can perform the fanout/fanin wire re-connections by modeling the problem of the wire reconnections by a flow graph and then solving the problem using the maxflow-mincut algorithm on the flow graph. However, the flow graph cannot fully characterize the wire re-connections which causes GFO to lose optimality on several obvious cases. In addition, we find that the fanin re-connection can have more optimization power than the fanout re-connection but requires more sophisticated modeling. In this paper we re-formulate the problem of the fanout/fanin re-connections by a new graph called the implication flow graph. We show that the problem of wire re-connections on the implication flow graph is NP complete and also propose an efficient heuristic on the new graph. Our experimental results are very exciting. Shih-Chieh Chang 0001, Zhong-Zhen Wu, He-Zhe Yu |
ICCAD | 1 |
| 2000 | Synthesis of CMOS Domino Circuits for Charge Sharing AlleviationabstractThe Charge Sharing (CS) problem is one of notorious noise problems in domino circuits design and test. In this paper, this problem is thoroughly investigated by considering circuit topology and circuit function. The sensitivity of each domino gate to the CS problem is represented by the concept of CS-vulnerability. A method to derive the CS-vulnerability and the test pattern for each domino gate is suggested. We also propose a transition reordering method to dramatically reduce the CS-vulnerabilities for all domino gates, so that the CS problem can be alleviated. Simulation results demonstrate that our transistor reordering method can efficiently reduce the CS-vulnerabilities for most of domino circuits. Ching-Hwa Cheng, Shih-Chieh Chang 0001, Shin-De Li, Wen-Ben Jone, Jinn-Shyan Wang |
ICCAD | 2 |
| 2000 | A timing-driven pseudo-exhaustive testing of VLSI circuitsabstractThe object of this paper is to reduce the delay penalty of bypass storage cell (bsc) insertion for pseudo-exhaustive testing. We first propose a tight delay lower bound algorithm which estimates the minimum circuit delay for each node after bsc insertion. By understanding how the lower bound algorithm loses optimality, we can propose a bsc insertion heuristic which tries to insert bscs so that the final delay is as close to the lower bound as possible. Our experiments show that the results of our heuristic are either optimal because they are the same as the delay lower bounds or they are very close to the optimal solutions. Shih-Chieh Chang 0001, Jiann-Chyi Rau |
ISCAS | 1 |
| 2000 | A compact factored form for a Boolean functionabstractA factored form of a Boolean function is a common representation to express the complexity of a Boolean function in multi-level logic. However, a factored form which inhibits the appearance of the inversion operation is still a restricted way of representing a multi-level circuit. In this paper, we present a novel representation of a Boolean function, called the invert factored form representation. This representation mainly takes advantage of the inversion of whole or part of a Boolean function so that fewer literals and better multi-level circuit implementation can be obtained. Based on this novel presentation, our algorithm attempts to find a minimal expression. Experimental results also show the literal counts based on the novel representation are smaller than those on the traditional factored form representation. Jiann-Chyi Rau, Yan-Min Chen, Shih-Chieh Chang 0001 |
ISCAS | 3 |
| 2000 | TAIR: testability analysis by implication reasoningabstractTo predict the difficulty of testing a wire stuck-at fault, testability analysis algorithms provide an estimated testability value by computing controllability and observability. In most common previous work such as COP and SCOAP, signal correlation between controllability and observability is not well handled. As a result, the estimated values can be quite inaccurate, On the other hand, some previous work can take into account signal correlation but may require more CPU time. This paper discusses an efficient method for testability analysis improvement. Our algorithm starts with results obtained from conventional testability analysis such as COP. For each stuck- at fault, we gradually refine these results by recursively applying some simple signal correlation rules. Experimental results show that, with reasonable run-time overhead, significant improvement for testability analysis can be achieved. Shih-Chieh Chang 0001, Wen-Ben Jone, Shi-Sen Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1999 | Gate-Level Design Exploiting Dual Supply Voltages for Power-Driven ApplicationsabstractThe advent of portable and high-density devices has made power consumption a critical design concern.In this paper, we address the problem of reducing power consumption via gate-level voltage scaling for those designs that are not under the strictest timing budget.We first use a maximum-weighted independent set formulation for voltage reduction on non-critical part of the circuit.Then, we use a minimum-weighted separator set formulation to do gate sizing and integrate the sizing procedure with a voltage scaling procedure to enhance power saving on the whole circuit.The proposed methods are evaluated using the MCNC benchmark circuits.and an average of 19.12% power reduction over the circuits having only one supply voltage has been achieved. Chingwei Yeh, Min-Cheng Chang, Shih-Chieh Chang 0001, Wen-Ben Jone |
DAC | 3 |
| 1999 | Synthesis for multiple input wires replacement of a gate for wiring considerationabstractThe alternative wire technique attempts to replace a target wire by another wire without changing the logic functionality. In this paper we propose two new transformations of replacing wires. One transformation simultaneously replaces multiple input wires of a gate by a new set of input wires and the other performs gate decomposition during the alternative wire process. To accomplish such complex transformations, we discuss some theoretical foundations for replacing multiple wires. Understanding how wires/gates can be replaced by other wires/gates allows us to speedup the process tremendously. Shih-Chieh Chang 0001, Jung-Cheng Chuang, Zhong-Zhen Wu |
ICCAD | 1 |
| 1999 | Circuit Optimization by RewiringabstractPresents a very efficient optimization method suitable for multi-level combinational circuits. The optimization is based on incremental restructuring of a circuit through a sequence of additions and removals of redundant wires. Our algorithm applies the techniques of automatic test pattern generation (ATPG), which can efficiently detect redundancies. During the ATPG process, certain nodes in the circuit must have particular logic assignments for a test to exist. Based on the properties of these mandatory assignments, we have developed theorems to eliminate unnecessary wire redundancy checking. This results in a significant performance improvement. The fast run time and the excellent scaling to large circuits make our Boolean optimization method practical for industrial applications. Shih-Chieh Chang 0001, Lukas P. P. P. van Ginneken, Malgorzata Marek-Sadowska |
IEEE Trans. Computers | 1 |
| 1999 | Efficient Boolean division and substitution using redundancy addition and removingabstractBoolean division, and hence Boolean substitution, produces better result than algebraic division and substitution. However, due to the lack of an efficient Boolean division algorithm, Boolean substitution has rarely been used. We present an efficient Boolean division and Boolean substitution algorithm. Our technique is based on the philosophy of redundancy addition and removal. By adding multiple wires/gates in a specialized way, we tailor the philosophy onto the Boolean division and substitution problem. From the viewpoint of traditional division/substitution, our algorithm can perform substitution not only in sum-of-product form but also in product-of-sum form. Our algorithm can also naturally take all types of internal don't cares into consideration. As far as substitution is concerned, we also discuss the case where we are allowed to decompose not only the dividend but also the divisor. Experiments are presented and the result is promising. Shih-Chieh Chang 0001, David Ihsin Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1998 | Efficient Boolean Division and SubstitutionabstractBo ole andivision, and henc eBo ole ansubstitution, produc es better result than algebraic division and substitution. However, due to the lack of an efficient Bo ole andivision algorithm, Bo ole ansubstitution has rarely b een used. We present an efficient Bo ole andivision and substitution algorithm. Our technique is based on the philosophy of redundancy addition and removal. By adding multiple wires/gates in a specialized way, we tailor the philosophy onto the Bo ole an division and substitution problem. F rom the viewpoint of traditional division/substitution, our algorithm can perform substitution not only in sum-of-product form for but also in product-of-sum form. Our algorithm can also naturally take all typ es of don't cares into consideration. As far as substitution is conc erne d, we also discuss the case where we are allowed to decompose not only the dividend but also the divisor. Experiments are presente d and the result is pr omising. Shih-Chieh Chang 0001, David Ihsin Cheng |
DAC | 1 |
| 1998 | A novel combinational testability analysis by considering signal correlationabstractTo predict the difficulty of testing a wire stuck-at fault, testability analysis algorithms provide an estimated testability value by computing controllability and observability. In all previous work, signal correlation between controllability and observability is generally ignored. As a result, the estimated value can be inaccurate. This paper discusses an efficient method to take into account signal correlation for testability analysis. Our experimental results have shown that, with little run time overhead, significant improvement of testability analysis can be achieved. Shih-Chieh Chang 0001, Shi-Sen Chang, Wen-Ben Jone, Chien-Chung Tsai |
ITC | 1 |
| 1998 | A tree-structured LFSR synthesis scheme for pseudo-exhaustive testing of VLSI circuitsabstractThis paper presents a new test architecture, called Tree-LFSR/SR, to more effectively generate pseudo-exhaustive test patterns for combinational VLSI circuits. Instead of using a single scan chain, the proposed test architecture routes a scan tree driven by the LFSR to generate all possible input patterns for each output cone. The new test architecture is able to take advantages of both signal sharing and signal reuse. The benefits are: (1) the hardware overhead can be greatly reduced by saving routing area and XOR circuits, and (2) the difficulty of test architecture synthesis can be eased by accelerating the searching process of appropriate residues. The Tree-LFSR/SR configuration is then extended, if necessary, by adding XOR networks to deal with more complex input-output relations. An efficient method to directly synthesize the XOR network is also included. Experimental results obtained by simulating combinational benchmark circuits are very encouraging. Wen-Ben Jone, Jiann-Chyi Rau, Shih-Chieh Chang 0001, Yu-Liang Wu |
ITC | 3 |
| 1997 | Postlayout logic restructuring using alternative wiresabstractIn this paper, we propose a layout-driven synthesis approach for field programmable gate arrays (FPGA's). The approach attempts to identify alternative wires and alternative functions for wires that cannot be routed due to the limited routing resources in FPGA. The alternative wires (in the logic level) that can be routed through less congested areas substitute the unroutable wires without changing the circuit's functionality. Allowing the logic blocks to have alternative functions also increases the chance of successful routing. A redundancy addition and removal technique is used to identify such alternative wires. Experimental results are presented to demonstrate the usefulness of this approach. For a set of randomly selected benchmark circuits, on the average, 30-50% of wires have alternative wires. These results indicate that the routing flexibility can be substantially increased by considering these alternative wires. Our prototype system successfully completed routing for two AT&T designs that cannot be handled by an FPGA router alone. The proposed synthesis technique can also be applied to standard cell and gate array designs to reduce the routing area. Shih-Chieh Chang 0001, Kwang-Ting Cheng, Nam Sung Woo, Malgorzata Marek-Sadowska |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1996 | Fast Boolean optimization by rewiringabstractThis paper presents a very efficient Boolean logic optimization method. The boolean optimization is achieved by adding and removing redundant wires in a circuit. Our algorithm applies the reasoning of Automatic Test Pattern Generation (ATPG) which can detect redundancy efficiently. During the ATPG process, mandatory assignments are assignments which must be satisfied. Our algorithm analyzes different characteristics of mandatory assignments during the ATPG process. New theoretical results based on the analysis are presented which lead to significant performance improvements. The fast run time and the excellent scaling to large problems make our Boolean optimization method practical for industrial applications. Experiments show that the optimization results are comparable to those of Kunz and Pradhan (1994) while the run time is two orders of magnitude faster (average 126/spl times/ speed up). Furthermore, we report optimization results for several large examples, which were previously thought to be too large to be handled by Boolean optimization methods. Shih-Chieh Chang 0001, Lukas P. P. P. van Ginneken, Malgorzata Marek-Sadowska |
ICCAD | 1 |
| 1996 | Perturb and simplify: multilevel Boolean network optimizerabstractIn this paper, we present logic optimization techniques for multilevel combinational networks. Our techniques apply a sequence of perturbations which result in simplification of the circuit. The perturbation and simplification is achieved through wires/gates addition and removal which are guided by the Automatic Test Pattern Generation (ATPG) based reasoning. The main operations of our approaches are incremental transformations of the circuit (such as adding wires/gates and changing gate's functionality) to remove some particular wire, At each iteration, a summary information of such wires/gates addition and removal is precomputed first. Then, a transformation is chosen to remove several wires at once. We have performed experiments on MCNC benchmarks and compared the results to those of misII and RAMBO. Experimental results are very encouraging. Shih-Chieh Chang 0001, Malgorzata Marek-Sadowska, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1996 | Technology mapping for TLU FPGAs based on decomposition of binary decision diagramsabstractThis paper proposes an efficient algorithm for technology mapping targeting table look-up (TLU) blocks. It is capable of minimizing either the number of TLUs used or the depth of the produced circuit. Our approach consists of two steps. First a network of super nodes, is created. Next a Boolean function of each super node with an appropriate don't care set is decomposed into a network of TLUs. To minimize the circuit's depth, several rules are applied on the critical portion of the mapped circuit. Shih-Chieh Chang 0001, Malgorzata Marek-Sadowska, TingTing Hwang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1995 | An Efficient Algorithm for Local Don't Care Sets CalculationabstractLocal don't cares of an internal node expressed in terms of its immediate inputs are usually of interest. One can directly apply any two-level minimizer on the on-set and the local don't cares set to simplify an internal node. In this paper, we propose a memory efficient technique to calculate local don't cares of internal nodes in a combinational circuit. Our technique of calculating local don't cares makes use of automatic test pattern generation (ATPG) approach which allows us to identify quickly whether a cube in the local space is a don't care or not. Unlike other approaches which construct an intermediate form of don't cares in terms of the primary inputs, our technique directly computes the don't care cubes in the local space. This gives us a significant advantage over the previous approaches in memory usage. Experimental results on MCNC benchmarks are very encouraging. Shih-Chieh Chang 0001, Malgorzata Marek-Sadowska, Kwang-Ting Cheng |
DAC | 1 |
| 1995 | Logic Synthesis for Engineering ChangeabstractIn the process of VLSI design, specifications are often changed. It is desirable that such changes will not lead to a very different design so that a large part of engineering effort can be preserved. We consider synthesis algorithms for handling such engineering changes. Given a synthesized network, our algorithm modifies it minimally to realize a new specification. Chih-Chang Lin, Kuang-Chien Chen, Shih-Chieh Chang 0001, Malgorzata Marek-Sadowska, Kwang-Ting Cheng |
DAC | 3 |
| 1994 | Layout Driven Logic Synthesis for FPGAsabstractIn this paper, we propose a layout driven synthesis approach for Field Programmable Gate Arrays (FPGAs). The approach attempts to identify alternative wires and alternative functions for wires that cannot be routed due to the limited routing resources in FPGA. The alternative wires (in the logic level) that can be routed through less congested areas substitute the unroutable wires without changing the circuit's functionality. Allowing the logic blocks to have alternative functions also increases the flexibility of routing. The redundancy addition and removal techniques are used to identify such alternative wires. Experimental results are presented to demonstrate the usefulness of this approach. For a set of randomly selected benchmark circuits, on the average, 30%-50% of wires have alternative wires. These results indicate that the routing flexibility can be substantially increased by considering these alternative wires. Our prototype system successfully completed the routing for two AT&T designs that cannot be handled by an FPGA router alone. The proposed synthesis technique can also be applied to standard cell and gate array designs to reduce the routing area. Shih-Chieh Chang 0001, Kwang-Ting Cheng, Nam Sung Woo, Malgorzata Marek-Sadowska |
DAC | 1 |
| 1994 | Perturb and simplify: multi-level boolean network optimizer
Shih-Chieh Chang 0001, Malgorzata Marek-Sadowska |
ICCAD | 1 |
| 1992 | Technology Mapping via Transformations of Function GraphsabstractThe authors address the problem of how to realize a given combinational circuit described by means of Boolean equations using the minimum number of blocks of the target TLU table lookup architecture. Their Boolean decomposition scheme works directly on a reduced ordered binary decision diagram (ROBDD) of a subject function, using two techniques. The first, referred to as cutting, is an efficient implementation of Roth-Karp decomposition. The second technique is referred to as a substitution. The idea is to replace subgraphs of ROBDD by new variables. The substitution process is accompanied by certain reductions of the resulting ROBDD graph, which further decreases its size.> Shih-Chieh Chang 0001, Malgorzata Marek-Sadowska |
ICCD | 1 |