EDBT 2026 Demo / reviewers in the wild / expert
Yu-Guang Chen
dblp:55/3324
· DBLP profile ↗
32ranked-venue papers
17as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 17 first-author · 15 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FastGDBN: A GPU-Accelerated DNN for Identifying Good Dies in Bad Neighborhoods
Yu-Heng Tsao, Yu-Guang Chen |
ETS | 2 |
| 2025 | Overcoming Training Data Scarcity in Routing Demand Prediction via Ensemble LearningabstractAs CMOS technology scales down, the number of standard cells increases rapidly. The increasing cell count raises the complexity of physical design. Routing is one of the most time-consuming stages in the physical design flow. When routing fails to meet design rules or performance targets, designers must revise earlier stages such as floorplanning or placement. Repeating the routing process causes high design cost and long time-to-market. Early prediction of routing demand helps reduce design iterations. An ensemble learning model based on XGBoost is proposed to predict global routing demand using placement-stage features. The XGBoost-based model achieves higher accuracy than CNN- and FCN-based models, improving R² by 0.12 and 0.125, respectively. The inference speed is also significantly faster, up to 14.95×. Feature importance analysis enables reduction of training and inference overhead with minimal accuracy loss. Yu-Guang Chen, Shih-Cheng Huang, Cheng-Hong Tsai, De-Shiun Fu, Mango Chia-Tso Chao |
ACM Great Lakes Symposium on VLSI | 1 |
| 2025 | Optimizing Reliability and Energy Efficiency in Heterogeneous Multicore Systems: A Novel Task Deployment StrategyabstractThe rapid advancement of CMOS technology has empowered modern integrated circuits (ICs) to handle increasingly complex tasks, which are crucial across various applications. However, achieving a balance between performance and energy consumption remains a challenge, particularly for edge devices such as cell phones and IoT sensors. Heterogeneous multicore systems, combining high-performance and energy-efficient cores, offer a promising solution for optimizing power while meeting performance requirements. Nevertheless, the continuous evolution of technology exacerbates reliability concerns in multicore systems, notably transient errors and aging effects. Previous studies have proposed various strategies, including task replication and Dynamic Voltage and Frequency Scaling (DVFS), to mitigate reliability degradation. However, the lack of holistic consideration for these mitigation techniques may inadvertently shorten the system's lifespan due to potential conflicts between them. Hence, this paper thoroughly analyzes the interplay between strategies for mitigating transient errors and aging effects. Subsequently, we introduce a novel reliability-aware task deployment framework aimed at significantly extending the system lifespan while adhering to a specified reliability target and minimizing overall energy consumption. Experimental findings demonstrate the efficacy of our framework, showcasing a 4.98x improvement in system lifespan and a 50% reduction in energy consumption compared to prior studies. Yu-Guang Chen, Yin-Rong Zhuo, Zheng-Wei Chen, Ing-Chao Lin |
ACM Great Lakes Symposium on VLSI | 1 |
| 2025 | Invited Paper: 2025 ICCAD CAD Contest Problem A: Hardware Trojan Detection on Gate-Level NetlistabstractThe increasing reliance on third-party intellectual property (IP) cores in modern integrated circuit (IC) design has introduced significant security vulnerabilities, particularly the risk of Hardware Trojans (HTs). These malicious modifications can compromise system integrity, leading to data leakage, unauthorized access, or functional failures. Traditional detection methods often depend on the availability of a golden chip, which is not always feasible. This paper presents the 2025 ICCAD CAD Contest Problem A, which challenges participants to develop machine learning-based solutions for detecting HTs directly from gate-level netlists without requiring a golden reference. The problem is formulated with defined Trojan behaviors, input/output specifications, and evaluation metrics, including correctness and F1 score. The contest aims to foster innovation in HT detection by leveraging advanced data-driven techniques and scalable analysis frameworks. Chung-Han Chou, Chih-Jen Hsu, Hung-Chun Chiu, Kai-Chiang Wu, Yu-Guang Chen, Zhuo Li 0001 |
ICCAD | 5 |
| 2025 | COPA: A Congestion-Oriented Pin Assignment Framework for Intra-Block Physical Design OptimizationabstractIn hierarchical integrated circuit (IC) designs, pin assignment plays a critical role in determining routing quality and overall layout efficiency. While conventional approaches primarily focus on wirelength minimization, the impact of feedthrough pin assignment on internal routing congestion has been relatively overlooked. Our observation reveals that proper assignment of feedthrough pins can significantly reduce routing congestion without increasing wirelength. In this paper, we present COPA, a Congestion-Oriented Pin Assignment framework that explicitly targets congestion mitigation within intra-block routing regions. By formulating the feedthrough pin assignment problem as a nonlinear optimization task, COPA employs a gradient-based algorithm that iteratively adjusts pin positions based on congestion severity, quantified through a routing-demand-driven metric. Experimental results on industrial testcases demonstrate that COPA achieves an average 7.6% reduction in total overflow, with a maximum improvement of up to 9%, compared to initial feedthrough pin assignments generated by a commercial EDA tool. Shu-Yi Tsai, Yu-Guang Chen, Kun-Min Chen, Sheng-Bing Ke, Chung-Hui Hsieh, Mango Chia-Tso Chao |
ICCAD | 2 |
| 2025 | Basis Sharing: Cross-Layer Parameter Sharing for Large Language Model CompressionabstractLarge Language Models (LLMs) have achieved remarkable breakthroughs. However, the huge number of parameters in LLMs require significant amount of memory storage in inference, which prevents their practical deployment in many applications. To reduce memory storage of LLMs, singular value decomposition (SVD) provides a promising solution to approximate weight matrices for compressing LLMs. In this paper, we take a step further to explore parameter sharing across different layers with SVD to achieve more effective compression for LLMs. Specifically, weight matrices in different layers are decomposed and represented with a linear combination of a set of shared basis vectors and unique coefficients. The types of weight matrices and the layer selection for basis sharing are examined when compressing LLMs to maintain the performance. Comprehensive experiments demonstrate that Basis-Sharing outperforms state-of-the-art SVD-based compression approaches, especially at large compression ratios. Jingcun Wang, Yu-Guang Chen, Ing-Chao Lin, Bing Li 0005, Grace Li Zhang |
ICLR | 2 |
| 2025 | DriverOmicsNet: an integrated graph convolutional network for multi-omics exploration of cancer driver genesabstractCancer is a complex and heterogeneous group of diseases driven by genetic mutations and molecular changes. Identifying and characterizing cancer driver gene is crucial for understanding cancer biology and guiding precision oncology. Integrating multi-omics data can reveal the intricate molecular interactions underlying cancer progression and treatment responses. We developed a graph convolutional network (GCN) framework, DriverOmicsNet, that integrates multi-omics data using STRING protein-protein interaction networks and correlation-based weighted gene correlation network analysis (WGCNA). We applied this framework to 15 cancer types, analyzing 5555 tumor samples to predict cancer-related features such as homologous recombination deficiency, cancer stemness, immune clusters, tumor stage, and survival outcomes. DriverOmicsNet demonstrated superior predictive accuracy and model performance metrics across all target labels when compared with GCN models based on STRING network alone. Gene expression emerged as the most significant feature, reflecting the dynamic and functional state of cancer cells. The combined use of STRING PPI and WGCNA networks enhanced the identification of key driver genes and their interactions. Our study highlights the effectiveness of using GCNs to integrate multi-omics data for precision oncology. The integration of STRING PPI and WGCNA networks provides a comprehensive framework that improves predictive power and facilitates the understanding of cancer biology, paving the way for more tailored treatments. Yang-Hong Dai, Po-Chien Shen, Wun-Long Jheng, Ding-Jie Lee, Yu-Guang Chen |
Briefings Bioinform. | 6 |
| 2024 | GNN-Based INC and IVC Co-Optimization for Aging MitigationabstractAs semiconductor processes advance, circuit aging becomes prominent. One of the most severe aging effects is Negative Bias Temperature Instability (NBTI), which increases the threshold voltage and the propagation delay of PMOS transistors. To mitigate NBTI, aging mitigation methods such as Internal Node Control (INC) and Input Vector Control (IVC) have been proposed. INC applies designed logic gates, while IVC uses appropriate input patterns during circuit idle. However, INC leads to extra area overhead and power consumption, and the circuit structure limits the controllability of IVC. Although various approaches have proposed aging tolerance methods with INC or IVC, only a few of them consider co-optimization. In this paper, we introduce a GNN-based INC and IVC co-optimization framework to minimize aging-induced delay. The key concept of our framework is using GNN to identify serious-aged gates in a circuit, and then using INC and IVC to mitigate the aging effect under an area overhead constraint. The experimental results indicate that our method reduces aging-induced delay and area by 2.16 times and 29.5%, respectively, compared to previous work. Yu-Guang Chen, Hsiu-Yi Yang, Ing-Chao Lin |
ETS | 1 |
| 2024 | Enhancing Stability in CRPs: A Novel Parallel Scan-Chain PUF Design Considering Aging EffectsabstractHardware Trojans and side-channel attacks pose significant threats such as sensitive information leaking and/or malfunctioning to modern cryptoprocessors. To overcome those threats, Physical Unclonable Function (PUF) has been considered one of the security primitives for secret keys or unique IDs due to the characteristic of stable, random, and unique responses. However, most existing PUF designs, which operate independently of the original circuit, are vulnerable to removal attacks and result in substantial resource overhead. To address these issues, authors in [5] proposed parallel scan-chain PUF, which is built on the standard Design-for-Testability (DFT) structure of scan flip-flops. However, aging effects, such as Bias Temperature Instability (BTI) and Hot Carriers Injection (HCI), significantly impact input conditions between two Scan Flip-Flops (SFFs) in normal mode, leading to complex Error Correction Codes (ECC) and increased area overhead. Therefore, we introduce a novel design to mitigate aging effects, enhance reliability, and analyze the parallel scan-chain PUF structure to identify error causes. Besides, to decrease the errors caused by aging effects, we use the proposed aging compensator to offset the delay between two different SFFs after aging mutually. Additionally, we incorporate signal gating for the arbiter, mitigating unbalanced aging effects and saving power consumption. Experimental results demonstrate that the proposed method can reduce the responses from an average error rate of 39.96% to less than 7.5% within 10 years. Moreover, with the complexity of ECC reduced, it offers ~8x overhead reduction for the Bose–Chaudhuri–Hocquenghem (BCH) encoder and decoder. Yu-Guang Chen, Tzong-Ying Lee |
ISCAS | 1 |
| 2024 | Aging-Aware Energy-Efficient Task Deployment of Heterogeneous Multicore SystemsabstractHeterogeneous multicore systems, which consist of high-performance and power-efficient cores, are emerging to satisfy the various demands on performance and power consumption. On the other hand, as CMOS technology continues to shrink in size, the aging effect, which can cause performance degradation or timing failures, has become a non-negligible threat to lifetime reliability. To overcome the challenges under the aging effect, various approaches have been proposed in previous studies. Most previous studies, however, did not consider the different characteristics of big and little cores. In addition, most of them do not consider critical tasks with the strict timing requirements present in real-time applications, resulting in early system failure. Therefore, considering different characteristics of cores and the presence of critical tasks, we propose an aging-aware task deployment framework for real-time systems. In this framework, for high-performance big cores, we propose a novel asymmetric aging-aware strategy. This strategy finds an energy-efficient task-to-core assignment to reserve some healthy cores at the early system life stage. The reserved cores are kept idle with the lowest voltage and can execute critical tasks at the late system life stage, extending the system lifetime. Meanwhile, the non-reserved cores use lower voltages to execute tasks, reducing the aging effect. For energy-efficient little cores, we adopt the symmetric aging-aware strategy to balance out the aging effect of each little core. With a balanced aging effect, the utilization of little cores is improved. In addition, we propose voltage/frequency boosting and task migration techniques to increase the number of cores that can meet the task timing constraints. Compared to the state of the art, the proposed framework can achieve 1.10x lifetime improvement and 5% energy reduction. Yu-Guang Chen, Chieh-Shih Wang, Ing-Chao Lin, Zheng-Wei Chen, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | An On-Line Aging Detection and Tolerance Framework for Improving Reliability of STT-MRAMsabstractSpin-transfer-torque magnetic random-access memory (STT-MRAM) is one of the most promising emerging memories for on-chip memory. However, the magnetic tunnel junction (MTJ) in the STT-MRAM suffers from several reliability threats which degrade the endurance, create defects, and cause memory failure. One of the primary reliability issues comes from time-dependent dielectric breakdown (TDDB) on MTJ, which deviates resistance value of MTJ over time and may lead to reading error. To overcome this challenge, in this paper we present an on-line aging detection and tolerance framework to dynamically monitor the electrical parameter deviations and provide appropriate compensation to avoid reading error. The on-line aging detection mechanism can identify aged words by monitoring read current and then the aging tolerance mechanism can adjust the reference resistance of the sensing amplifier to compensate the aging-induced resistance drop of MTJ. In comparison with existing testing-based aging detection techniques, our mechanism can operate on-line with read operations for both aging detection and tolerance simultaneously with negligible performance overhead. Simulation and analysis results show that the proposed techniques can successfully detect 99% aging words under process variation and achieve at most 25% reliability improvement of STT-MRAMs. Yu-Guang Chen, Po-Yeh Huang, Jin-Fu Li 0001 |
ASP-DAC | 1 |
| 2023 | DRC Violation Prediction with Pre-global-routing Features Through Convolutional Neural NetworkabstractDesign Rule Checking (DRC) is one of the most important metrices in physical design procedure to evaluate quality of a detail route. The prediction of DRC violation (DRV) in the early stage can reduce the iterations of design procedure and improve the efficiency of the physical design closure. Several researchers have applied machine-learning techniques to predict the DRVs of a detail route at different design stages with various input features. In this paper, we proposed a machine learning model to predict DRVs with the information obtained after placement stage. Specifically, we build a ResNet-like CNN model to predict whether a DRV may occur in a targeted grid after detail route. Our features consist of not only quantified placement information but also layout-image features to take pin accessibility into account for better prediction result. Moreover, we apply an under-sampling technique to select critical training samples to improve the training efficiency. A series of experiments have been conducted and the results show that compared with previous works, our prediction result can outperform Fully Convolutional Network (FCN) based approaches. Jhen-Gang Lin, Yu-Guang Chen, Yun-Wei Yang, Wei-Tse Hung, Cheng-Hong Tsai, De-Shiun Fu, Mango Chia-Tso Chao |
ACM Great Lakes Symposium on VLSI | 2 |
| 2023 | Invited Paper: Overview of 2023 CAD Contest at ICCADabstractThe “CAD Contest at ICCAD” is a challenging, multi-month, research and development competition, focusing on advanced, real-world problems in the field of electronic design automation (EDA). Since 2012, the contest has been publishing many sophisticated circuit design problems, from system-level design to physical design, together with industrial benchmarks and solution evaluators. Contestants can participate in one or more problems provided by EDA/IC industry. The winners will be awarded at an ICCAD special session dedicated to this contest. Every year, the contest attracts more than a hundred teams, fosters productive industry-academia collaborations, and leads to hundreds of publications in top-tier conferences and journals. The 2023 CAD Contest has 210 teams from all over the world, which generates the highest participation record. Moreover, the problems of this year cover state-of-the-art EDA research trends such as circuit verification, hardware security, 3D-IC, and Machine Learning (ML) for EDA from well-known EDA/IC companies. We believe the contest keeps enhancing impact and boosting EDA researches. Takashi Sato 0001, Chun-Yao Wang, Yu-Guang Chen, Tsung-Wei Huang |
ICCAD | 3 |
| 2023 | DRC Violation Prediction After Global Route Through Convolutional Neural NetworkabstractDesign rule checking (DRC) violation (DRV) prediction with early stage design information can help to reduce the iterations of design procedure and can speed up the physical-design closure. It is known that accurately predicting detailed routing-level DRV with information obtained at global route (GR) stage can significantly speed up the design closure. However, without sufficient prediction accuracy, the result may lead to suboptimal design or even longer design time. Therefore, in this article, we propose two machine-learning frameworks to predict the detailed routing-level DRV map of a given design. The first framework is based on the congestion report obtained at global routing stage, and the second framework considers both the placement information and the congestion report of global routing. We then compare the runtime and accuracy of the two models. The proposed frameworks utilize convolutional neural network as the core technique to train these prediction models. The training dataset is collected from 15 industrial designs using a leading commercial automatic placement and routing (APR) tool, and the total number of collected training samples exceeds 26M. A specialized under-sampling technique is also proposed to select important training samples for learning, compensate for the inaccuracy misled by a highly imbalanced training dataset, and speed up the entire training process. The experimental results demonstrate that our both models can result in not only a significantly higher accuracy than previous related works, but also a DRV map visually matching the actual ones closely. The average runtime of using our learned model from the first framework to generate a DRV map is only 3% of global routing, and the prediction accuracy of our learned model from the second framework can improve 7.6% compared to the one from the first framework. Our proposed framework can be viewed as a simple add-on tool to a current commercial placement and global router that can efficiently and effectively generate a more realistic DRV map without really applying detailed routing. Wei-Tse Hung, Yu-Guang Chen, Jhen-Gang Lin, Yun-Wei Yang, Cheng-Hong Tsai, Mango Chia-Tso Chao |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2022 | Overview of 2022 CAD Contest at ICCADabstractThe "CAD Contest at ICCAD" is a challenging, multi-month, research and development competition, focusing on advanced, real-world problems in the field of electronic design automation (EDA). Since 2012, the contest has been publishing many sophisticated circuit design problems, from system-level design to physical design, together with industrial benchmarks and solution evaluators. Contestants can participate in one or more problems provided by EDA/IC industry. The winners will be awarded at an ICCAD special session dedicated to this contest. Every year, the contest attracts more than a hundred teams, fosters productive industry-academia collaborations, and leads to hundreds of publications in top-tier conferences and journals. The 2022 CAD Contest has 166 teams from all over the world. Moreover, the problems of this year cover state-of-the-art EDA research trends such as circuit security, 3D-IC, and design space exploration from well-known EDA/IC companies. We believe the contest keeps enhancing impact and boosting EDA researches. Yu-Guang Chen, Chun-Yao Wang, Tsung-Wei Huang, Takashi Sato 0001 |
ICCAD | 1 |
| 2021 | Overview of 2021 CAD Contest at ICCADabstractThe “CAD Contest at ICCAD” is a challenging, multi-month, research and development competition, focusing on advanced, real-world problems in the field of electronic design automation (EDA). Since 2012, the contest has been publishing many sophisticated circuit design problems, from system-level design to physical design, together with industrial benchmarks and solution evaluators. Contestants can participate in one or more problems provided by EDA/IC industry. The winners will be awarded at an ICCAD special session dedicated to this contest. Every year, the contest attracts more than a hundred teams, fosters productive industry-academia collaborations, and leads to hundreds of publications in top-tier conferences and journals. The 2021 CAD Contest has 137 teams from all over the world. The contest keeps enhancing impact and boosting EDA research. Tsung-Wei Huang, Yu-Guang Chen, Chun-Yao Wang, Takashi Sato 0001 |
ICCAD | 2 |
| 2021 | An efficient NBTI-aware wake-up strategy: Concept, design, and manipulation
Yu-Guang Chen, Ing-Chao Lin, Kun-Wei Chiu, Cheng-Hsuan Liu |
Integr. | 1 |
| 2020 | Power Distribution Network Generation for Optimizing IR-Drop Aware TimingabstractAs supply voltage keeps scaling, the timing of an IC becomes more sensitive to the IR-drop of its power distribution network (PDN). In conventional timing signoff, designers assign a fixed voltage to all cells during the entire analysis process, which leads to over-design since the exact voltages at most cells are larger than the assigned value for timing analysis. To reflect the actual timing of a circuit, IR-drop aware timing analysis should be applied for timing signoff. This paper proposes a framework that can automatically generate a refined PDN for optimizing IR-drop aware timing based on a given initial PDN and cell placement. Our framework uses two novel timing-cost indexes to measure the impact of a PDN on IR-drop aware timing and a novel dynamic-programming based algorithm to obtain an optimal timing-driven PDN while efficiently handling the discontinuity of power rails caused by macros. The experiments based on various 28nm industrial designs demonstrate that our framework can always generate a PDN resulting in less worst-case negative slack, less total negative slack, and less timing-violating paths, compared to a previous work focusing on minimizing routing overhead. Our framework accesses all design information and makes all PDN changes through commercial EDA tools, and hence can be easily integrated into industrial design flows. Wen-Hsiang Chang, Li-Yi Lin, Yu-Guang Chen, Mango Chia-Tso Chao |
ICCAD | 3 |
| 2020 | Selective Sensor Placement for Cost-Effective Online Aging Monitoring and ResilienceabstractAggressive technology scaling trends, such as thinner gate oxide without proportional downscaling of supply voltage, aggravate the aging impact and thus necessitate an aging-aware reliability verification and optimization framework during early design stages. In this paper, we propose a novel in-situ sensing strategy based on deploying transition detectors (TDs), for on-chip aging monitoring and resilience. Transformed into the set cover problem and then formulated into maximum satisfiability, the proposed problem of TD/sensor placement can be solved efficiently. Experimental results show that, by introducing at most 2.2% area overhead (for TD/sensor placement), the aging behavior of a target circuit can be effectively monitored, and the correctness of its functionality can be perfectly guaranteed with an average of 77% aging resilience achieved. In other words, with 2.2% area overhead, potential aging-induced timing errors can be detected and then eliminated, while achieving 77% recovery from aging-induced performance degradation. Hao-Chun Chang, Li-An Huang, Kai-Chiang Wu, Yu-Guang Chen |
ISPD | 4 |
| 2019 | ROAD: Improving Reliability of Multi-core System via Asymmetric AgingabstractNegative-Bias Temperature Instability (NBTI), which may lead to performance degradation or even timing failure, has become one of the most drastic challenges in modern multi-core systems. To tolerate NBTI and extend the lifetime of the system, previous researchers proposed maintaining all cores in the multi-core system under similar aging conditions (symmetric aging) through various task assignment algorithms and/or dynamic voltage frequency scaling. Although the concept of symmetric aging provides efficient approaches to tolerating NBTI, it may reduce the lifetime of a multi-core system. If a critical task (i.e., a task with tight timing constraints) arrives when the system has already operated for years, it is possible that none of the equivalently aged cores will be able to complete the critical task within its timing constraints. This unavoidable timing failure then will shorten the lifetime of the system. In contrast, if a few cores are kept robust, these cores can be used to execute the critical task even if all the other cores are aged (asymmetric aging), which avoids timing failure and extends the system lifetime. Based on the above observation, this paper proposes a novel reliability improvement framework that consists of task graph Retiming, task Ordering, task Assignment under asymmetric aging, and Dynamic voltage selection (ROAD) for multi-core systems. With our framework, asymmetric aging can extend the system lifetime through successfully executing critical tasks at the later life stages of the system. The experimental results show that our approach can significantly increase the system lifetime with no or insignificant energy overhead. Yu-Guang Chen, Ing-Chao Lin, Jian-Ting Ke |
ICCAD | 1 |
| 2018 | An efficient NBTI-aware wake-up strategy for power-gated designsabstractThe wake-up process of a power-gated design may induce an excessive surge current and threaten the signal integrity. A proper wake-up sequence should be carefully designed to avoid surge current violations. On the other hand, PMOS sleep transistors may suffer from the negative-bias temperature instability (NBTI) effect which results in decreased driving current. Conventional wake-up sequence decision approaches do not consider the NBTI effect, which may result in a longer or unacceptable wake-up time after circuit aging. Therefore, in this paper, we propose a novel NBTI-aware wake-up strategy to reduce the average wake-up time within a circuit lifetime. Our strategy first finds a set of proper wake-up sequences for different aging scenarios (i.e. after a certain period of aging), and then dynamically reconfigures the wake-up sequences at runtime. The experimental results show that compared to a traditional fixed wake-up sequence approach, our strategy can reduce average wake-up time by as much as 45.04% with only 3.7% extra area overhead for the reconfiguration structure. Kun-Wei Chiu, Yu-Guang Chen, Ing-Chao Lin |
DATE | 2 |
| 2017 | Pattern based runtime voltage emergency prediction: An instruction-aware block sparse compressed sensing approachabstractThe relentless technology scaling calls for reduced supply voltage for dynamic power suppression. On the other hand, transistor threshold voltage cannot be scaled at the same pace to avoid excessive leakage power. Consequently, the noise margin is significantly reduced, leading to the deployment of various noise management systems that handle runtime voltage emergencies. Most of these systems rely on on-chip noise sensors, which are large in size and consume significant power. To tackle this issue, in this paper we propose a sensor-less voltage emergency estimation framework. It explores the relationship between switching activities and noise, and takes advantage of block sparse compressed sensing developed by the signal processing society. Experimental results on a few industrial designs show that by monitoring registers, voltage emergencies can be successfully predicted. Yu-Guang Chen, Michihiro Shintani, Takashi Sato 0001, Yiyu Shi 0001, Shih-Chieh Chang 0001 |
ASP-DAC | 1 |
| 2016 | A novel low-cost dynamic logic reconfigurable structure strategy for low power optimizationabstractLow power design techniques have been extensively applied in modern IC designs to avoid negative side effects from high power density. Unlike Dynamic Voltage and/or Frequency Scaling (DVFS) approaches only applied on a “fixed” design, we propose a dynamic logic reconfigurable structure strategy which allows dynamic switching from a high speed/power logic structure to a low speed/power logic structure. A design with such configurable structure is called Dynamic Logic Reconfigurable Structure (DLRS). Different from approximate computing which trades off between computation accuracy and power, our DLRS designs maintain data integrity. In this paper, we propose novel low-cost DLRS adders and multipliers, and a comprehensive framework for low power designs. We further integrate DLRS with DVFS, which creates more flexibility to trade-off between performance and power consumption. Experimental results show that with DLRS adders and multipliers in three indoor designs, the proposed method can achieve up to 60.05% power reduction compared with traditional DVFS scheme with only 6.55% area overhead. Yu-Guang Chen, Wan-Yu Wen, Yun-Ting Wang, You-Luen Lee, Shih-Chieh Chang 0001 |
ASP-DAC | 1 |
| 2015 | Q-Learning Based Dynamic Voltage Scaling for Designs with Graceful DegradationabstractDynamic voltage scaling (DVS) has been widely used to suppress power consumption in modern designs. The decision of optimal operating voltage at runtime should consider the variations in workload, process as well as environment. As these variations are hard to predict accurately at design time, various reinforcement learning based DVS schemes have been proposed in the literature. However, none of them can be readily applied to designs with graceful degradation, where timing errors are allowed with bounded probability to trade for further power reduction. In this paper, we propose a Q-learning based DVS scheme dedicated to the designs with graceful degradation. We compare it with two deterministic DVS schemes, i.e., a stepping based scheme and a statistical modeling based scheme. Experimental results on three 45nm industrial designs show that the proposed Q-learning based scheme can achieve up to 83.9% and 29.1% power reduction respectively with 0.01 timing error probability bound. To the best of the authors' knowledge, this is the first in-depth work to explore reinforcement learning based DVS schemes for designs with graceful degradation. Yu-Guang Chen, Wan-Yu Wen, Yiyu Shi 0001, Shih-Chieh Chang 0001 |
ISPD | 1 |
| 2015 | Novel Spare TSV Deployment for 3-D ICs Considering Yield and Timing ConstraintsabstractIn 3-D integrated circuits, through silicon via (TSV) is a critical enabling technique to provide vertical connections. However, it may suffer from many reliability issues such as undercut, misalignment, or random open defects. Various fault-tolerance mechanisms have been proposed in literature to improve yield, at the cost of significant area overhead. In this paper, we focus on the structure that uses one spare TSV for a group of original TSVs, and study the optimal assignment of spare TSVs under yield and timing constraints to minimize the total area overhead. We show that such problem can be modeled as a constrained graph decomposition problem. Two efficient heuristics are further developed to address this problem. Experimental results show that under the same yield and timing constraints, our heuristic can reduce the area overhead induced by the fault-tolerance mechanisms by up to 61%, compared with a seemingly more intuitive nearest-neighbor-based heuristic. Yu-Guang Chen, Wan-Yu Wen, Yiyu Shi 0001, Wing-Kai Hon, Shih-Chieh Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2014 | Critical Path Monitor Enabled Dynamic Voltage Scaling for Graceful Degradation in Sub-Threshold DesignsabstractSub-threshold designs play an important role in energy-constrained applications. In those designs, path delays depend exponentially on threshold voltage/temperature. As such, dynamic configurations at runtime are desired for best trade-off between operating power and performance. Unfortunately, most existing works only consider either process or temperature variations but not both, resulting in sub-optimal configurations or even functional failures. Moreover, little study has been performed on the graceful degradation of sub-threshold designs, which is important in the presence of drastic delay variations. Towards this, we present a novel critical path monitor based dynamic voltage scaling scheme. Considering both process and temperature variations, it minimizes the operating power under a given timing error probability (TEP) bound. An exact method to decide the optimal switching thresholds is also proposed. Experimental results on 45nm industrial designs show that with only 1% TEP, our scheme can reduce the operating power by up to 75.3% compared with the constant voltage scheme. To the best of the authors' knowledge, this is the very first work on dynamic configuration for graceful degradation in sub-threshold designs. Yu-Guang Chen, Kuan-Yu Lai, Wan-Yu Wen, Yiyu Shi 0001, Shih-Chieh Chang 0001 |
DAC | 1 |
| 2014 | Yield and timing constrained spare TSV assignment for three-dimensional integrated circuitsabstractThrough Silicon Via (TSV) is a critical enabling technique in three-dimensional integrated circuits (3D ICs). However, it may suffer from many reliability issues. Various fault-tolerance mechanisms have been proposed in literature to improve yield, at the cost of significant area overhead. In this paper, we focus on the structure that uses one spare TSV for a group of original TSVs, and study the optimal assignment of spare TSVs under yield and timing constraints to minimize the total area overhead. We show that such problem can be modeled through constrained graph decomposition. An efficient heuristic is further developed to address this problem. Experimental results show that under the same yield and timing constraints, our heuristic can reduce the area overhead induced by the fault-tolerance mechanisms by up to 38%, compared with a seemingly more intuitive nearest-neighbor based heuristic. Yu-Guang Chen, Kuan-Yu Lai, Ming-Chao Lee, Yiyu Shi 0001, Wing-Kai Hon, Shih-Chieh Chang 0001 |
DATE | 1 |
| 2014 | Multibit Retention Registers for Power Gated Designs: Concept, Design, and DeploymentabstractRetention registers have been widely used in power gated designs to store data during sleep mode. However, their excessive area and leakage power render it imperative to minimize the total retention storage size. The current industry practice replaces all registers with singlebit retention ones, which significantly limits the design freedom and yields suboptimal designs. Toward this, for the first time in the literature, we propose the concept and the design of multibit retention registers, with which only selected registers need to be replaced. The technique can significantly reduce the number of bits that need to be stored and thus the leakage power, but needs several clock cycles for mode transition. In addition, an efficient assignment algorithm is developed to minimize the total retention storage size subject to mode transition latency constraint. Experimental results show that our framework on average can reduce the leakage power in sleep mode by 84% along with additional mode transition latency of 6 to 11 clock cycles, compared with the singlebit retention register-based design. Yu-Guang Chen, Hui Geng, Kuan-Yu Lai, Yiyu Shi 0001, Shih-Chieh Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2012 | Efficient multiple-bit retention register assignment for power gated design: Concept and algorithmsabstractRetention registers have been widely used in power gated design to store data during sleep mode. Since they consume much larger area and power than normal registers, it is imperative to minimize the total retention storage size. The current industry practice only replace all registers with single-bit retention ones, which significantly limits the design freedom and results in excessive area and power overhead. Towards this, for the first time in literature, we propose the concept of multi-bit retention register, with which only selected registers need to be replaced. It can significantly reduce the number of bits that need to be stored and thus the area and leakage power, but needs several clock cycles for mode transition. In addition, an efficient assignment algorithm is developed to minimize the total retention storage size subject to mode transition latency constraint. Experimental results show that our framework on average can reduce the leakage power in sleep mode and the retention storage area by 66.03%, compared with the single-bit retention register based design. Yu-Guang Chen, Yiyu Shi 0001, Kuan-Yu Lai, Hui Geng, Shih-Chieh Chang 0001 |
ICCAD | 1 |
| 2012 | Efficient on-line module-level wake-up scheduling for high performance multi-module designsabstractPower consumption has become the major bottleneck for modern high-performance architectures, which typically contain large numbers of modules. To suppress leakage power, sleep transistors have been extensively used, and wake-up scheduling is needed to determine the wake-up times and order of these sleep transistors. Most existing works on wake-up scheduling are based on sleep transistors and delay buffers in daisy-chains; they work well for the gate-level scheduling within a module when all the gates need to be turned on. Yet, for state-of-the-art designs, the number of modules that need to be turned on and their locations may vary depending on the task to be performed at runtime. Accordingly, we cannot extend the existing gate-level scheduling algorithms to decide the module-level wake-up order. To address the problem, we propose to first off-line construct a multi-conflict graph (MCG) based on the noise constraints; based on the graph, we then develop an on-line algorithm to decide the wake-up order. Experimental results show that on average, the wake-up latency from our approach is not only 46.01% shorter compared with the existing work but also conservatively only 0.45% longer than that from a Monte Carlo search-based evaluation, which is orders of magnitude slower. To the best of our knowledge, this is the first in-depth study on on-line module-level wake-up scheduling for high-performance architectures. Ming-Chao Lee, Yiyu Shi 0001, Yu-Guang Chen, Diana Marculescu, Shih-Chieh Chang 0001 |
ISPD | 3 |
| 2011 | NBTI-aware power gating designabstractA header-based power gating structure inserts PMOS as sleep transistors between the power rail and the circuit. Since PMOS sleep transistors in the functional mode are turned-on continuously, Negative Bias Temperature Instability (NBTI) influences the lifetime reliability of PMOS sleep transistors seriously. To tolerate NBTI effect, sizes of PMOS sleep transistors are normally over-sized. In this paper, we propose a novel NBTI-aware power gating architecture to extend the lifetime of PMOS sleep transistors. In our structure, sleep transistors are switched on/off periodically so that overall turned-on times of sleep transistors are reduced and sleep transistors are less influenced by NBTI effect. The experimental results show that our approach can achieve better lifetime extensions of PMOS sleep transistors than previous works and few area overheads. Ming-Chao Lee, Yu-Guang Chen, Ding-Kei Huang, Shih-Chieh Chang 0001 |
ASP-DAC | 2 |
| 2006 | Tracers placement for IP traceback against DDoS attacksabstractThis paper explores the tracers deployment problem for IP traceback methods how many and where the tracers should be deployed in the network to be effective for locating the attack origins. The minimizing the number of tracers deployment problems depended on locating the attack origins are defined. The problem is proved to be NP-complete. A heuristic method which can guarantee that the distance between any attack origin and its first met tracer be within an assigned distance is proposed. The upper bound for the probability of an undetected attack node can be calculated in advance and used to evaluate the number of tracers needed for the proposed heuristic method. Extended simulations are performed to study the performance of the tracers deployment. Chun-Hsin Wang, James Chang Wu Yu, Chiu-Kuo Liang, Kun-Ming Yu, Wen Ouyang, Ching-Hsien Hsu, Yu-Guang Chen |
IWCMC | 7 |