Yu-Shih Su

dblp:27/6936 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 5 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Electronic design automation · 60% Hardware reliability and fault tolerance · 32% Integrated circuit design · 8%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
physical design
0.432013
Through-Silicon Via Fault-Tolerant Clock Networks for 3-D ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2013
Fault-tolerant 3D clock network · DAC 2011
Clock Skew Minimization in Multi-Voltage Mode Designs Using Adjustable Delay Buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010
Hardware reliability and fault tolerance
fault-tolerant design
0.322013
Through-Silicon Via Fault-Tolerant Clock Networks for 3-D ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2013
Fault-tolerant 3D clock network · DAC 2011
Hardware reliability and fault tolerance › fault-tolerant design
TSV fault tolerance
0.322013
Through-Silicon Via Fault-Tolerant Clock Networks for 3-D ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2013
Fault-tolerant 3D clock network · DAC 2011
Electronic design automation › physical design › clock network synthesis
clock tree synthesis
0.222011
Fault-tolerant 3D clock network · DAC 2011
Clock Skew Minimization in Multi-Voltage Mode Designs Using Adjustable Delay Buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010
Electronic design automation › physical design
clock network synthesis
0.212013
Through-Silicon Via Fault-Tolerant Clock Networks for 3-D ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2013
Electronic design automation › physical design › clock network synthesis
adjustable delay buffer allocation
0.112010
Clock Skew Minimization in Multi-Voltage Mode Designs Using Adjustable Delay Buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010
Electronic design automation › physical design › clock network synthesis
clock skew optimization
0.112010
Clock Skew Minimization in Multi-Voltage Mode Designs Using Adjustable Delay Buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010
Integrated circuit design
digital circuit design
0.112007
An Efficient Mechanism for Performance Optimization of Variable-Latency Designs · DAC 2007
Electronic design automation › physical design
timing optimization
0.112007
An Efficient Mechanism for Performance Optimization of Variable-Latency Designs · DAC 2007
Integrated circuit design
variable-latency design
0.112007
An Efficient Mechanism for Performance Optimization of Variable-Latency Designs · DAC 2007

Methods — techniques the papers use, named apart from their topics

redundant clock trees · 0.2TSV fault-tolerant unit · 0.2linear-time optimal assignment · 0.1heuristic buffer placement · 0.1telescopic units · 0.1short path activation · 0.1hold logic · 0.1
YearPublicationVenuePosition
2013 Benchmarking for research in power delivery networks of three-dimensional integrated circuits
abstract
Power integrity is generally considered to be one of the major bottlenecks hindering the prevalence of three-dimensional integrated circuits (3D ICs). The higher integration density and smaller footprint result in significantly increased power density, which threatens the system reliability. In view of this, there has been groundswell of interest in academia to model, design or optimize the power delivery networks (PDNs) in 3D ICs. Unfortunately, while several PDN benchmarks exist for 2D PDNs, none is available in the context of 3D. As a consequence, most existing literature resorts to ad-hoc designs by artificially stacking 2D PDNs for experiments, rendering the results less convincing. In this paper, we put forward a set of ten PDN benchmarks that are extracted from industrial 3D designs. These designs are carefully selected such that they cover a wide range of functionality, size, TSV number, tier number and packaging style. We hope that the released benchmarks can facilitate and promote research in 3D PDNs.
Pei-Wen Luo, Chun Zhang 0003, Yung-Tai Chang, Liang-Chia Cheng, Hung-Hsie Lee, Bih-Lan Sheu, Yu-Shih Su, Ding-Ming Kwai, Yiyu Shi 0001
ISPD7
2013 Through-Silicon Via Fault-Tolerant Clock Networks for 3-D ICs
abstract
Clock network synthesis is one of the most important and challenging problems in 3-D ICs. The clock signals have to be delivered by through-silicon vias (TSVs) to different tiers with minimum skew. While there are a few related works in literature, none consider the reliability of TSVs in a clock tree. Accordingly, the failure of any TSV in the clock tree yields a bad chip. The naive solution using double-TSV can alleviate the problem, but the significant area overhead renders it less practical for large designs. In this paper, we propose a novel TSV fault-tolerant unit (TFU) to provide tolerance against TSV failures. The TFU makes use of the existing 2-D redundant trees designed for prebond testing, and thus has minimum area overhead. In addition, the number of TSVs in a TFU is also adjustable to allow flexibility during clock network synthesis. Compared with the conventional double TSV technique, the 3-D clock network constructed by TFUs can achieve 58% area overhead reduction with similar yield rate on an industrial case. To the best of the authors' knowledge, this is the first work in the literature that considers the fault tolerance of a 3-D clock network. It can be easily integrated with any bottom-up clock network synthesis algorithm.
Chiao-Ling Lung, Yu-Shih Su, Hsih-Hsiu Huang, Yiyu Shi 0001, Shih-Chieh Chang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2012 Capturing the phantom of the power grid - on the runtime adaptive techniques for noise reduction
abstract
Power supply noise has become one of the primary concerns in low power designs. To ensure power integrity, designers need to make sure that voltage droop and bounce do not exceed noise margin in all possible scenarios. Since it is very difficult to capture the exact worst corner among the mist of complex functionalities in modern VLSI designs, statistical design methodologies have been adapted, which may bring significant design overhead. In view of this, various runtime techniques have been proposed in literature to suppress power grid noise adaptively. This paper first presents various challenges in power grid designs from an industrial perspective, explains the difficulties in handling them at deign time, and then reviews various runtime techniques to adaptively suppress power supply noise, including sensor-based power gating, re-routable decaps, proactive clock frequency actuator, and PLL based clocking.
Pei-Wen Luo, Yu-Shih Su, Liang-Chia Cheng, Ding-Ming Kwai, Yiyu Shi 0001
ASP-DAC3
2011 Fault-tolerant 3D clock network
abstract
Clock tree synthesis is one of the most important and challenging problems in 3D ICs. The clock signals have to be delivered by through-silicon vias (TSVs) to different tiers with minimum skew and latency. While there are a few related works in literature, none of them considers the reliability of TSVs. Accordingly, the failure of any TSV in the clock tree yields a bad chip. The naive solution using double-TSV can alleviate the problem. But the significant area overhead renders it less practical for large designs. In this paper, we propose a novel TSV fault-tolerant unit (TFU) that can provide tolerance against TSV failures in a 3D clock network. It makes use of the existing 2D redundant trees designed for pre-bond testing, and thus has minimum area overhead. Compared to the double TSV technique, the 3D clock network constructed by our TFUs can achieve 61% area reduction with 3.9% yield rate improvement on an industrial case. To the best of the authors' knowledge, this is the first practical work in literature that considers the fault tolerance of a 3D clock network.
Chiao-Ling Lung, Yu-Shih Su, Shih-Hsiu Huang, Yiyu Shi 0001, Shih-Chieh Chang 0001
DAC2
2011 Performance Optimization Using Variable-Latency Design Style
abstract
In many designs, the worst-case delay of a critical path may be activated infrequently. Traditional optimization approaches assume the worst-case conditions, which could lead to an inefficient resource usage. It is possible to improve the throughput of such designs by introducing variable latency. One existing realization of the variable-latency design style is based on telescopic units. The design of the hold logic in telescopic units influences the circuit's throughput. In this paper, we show that the traditionally designed hold logic may be inaccurate. We use the short path activation conditions to obtain more accurate hold logic and improve the efficiency of telescopic units. To reduce the overhead for large circuits, we propose an efficient heuristic methodology of constructing non-exact hold logic. We also discuss how to choose the telescopic unit's timing constraint. On average, our approach achieves the performance gain of 21.67% compared to 13.99%, reported in the previous work.
Yu-Shih Su, Da-Chung Wang, Shih-Chieh Chang 0001, Malgorzata Marek-Sadowska
IEEE Trans. Very Large Scale Integr. Syst.1
2010 Clock Skew Minimization in Multi-Voltage Mode Designs Using Adjustable Delay Buffers
abstract
In synchronous circuit designs, clock skew is difficult to minimize because a single physical layout of a clock tree must satisfy multiple constraints in a complicated power mode environment where certain modules may operate with different voltages. In this paper, we use adjustable delay buffers (ADB) whose delays can be tuned or adjusted to minimize clock skew under different power modes. Assuming that the positions of$k$ADBs are already determined, we first propose a linear-time optimal algorithm which assigns the values of ADBs so that the skew is optimal among all possible ADB assignments with a possibility of latency penalty. Then, we propose a modified optimal algorithm without latency penalty. We also propose an efficient heuristic to determine good positions for ADBs. Our results show significant improvement when compared to cases without ADBs.
Yu-Shih Su, Wing-Kai Hon, Cheng-Chih Yang, Shih-Chieh Chang 0001, Yeong-Jar Chang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2009 Value assignment of adjustable delay buffers for clock skew minimization in multi-voltage mode designs
abstract
In synchronous circuit designs, clock skew is difficult to minimize because a single physical layout of a clock tree must satisfy multiple constraints in a complicated power mode environment where certain modules may operate with different voltages. In this paper, we use Adjustable Delay Buffers (ADB) whose delays can be tuned or adjusted to minimize clock skew under different power modes. Assuming that the positions of k ADBs are already determined, we propose a linear-time optimal algorithm which assigns the values of ADBs so that the skew is optimal among all possible ADB assignments. We also propose an efficient heuristic to determine good positions for ADBs. Our results show significant improvement when compared to cases without ADBs. Categories and Subject Descriptors B.6.3 [Logic Design]: Design Aids — Optimization General Terms: Algorithms, Reliability
Yu-Shih Su, Wing-Kai Hon, Cheng-Chih Yang, Shih-Chieh Chang 0001, Yeong-Jar Chang
ICCAD1
2008 Synthesis of a novel timing-error detection architecture
abstract
Delay variation can cause a design to fail its timing specification. Ernst et al. [2003] observe that the worst delay of a design is least probable to occur. They propose a mechanism to detect and correct occasional errors while the design can be optimized for the common cases. Their experimental results show significant performance (or power) gain as compared with the worst-case design. However, the architecture in Ernst et al. [2003] suffers the short path problem, which is difficult to resolve. In this article, we propose a novel error-detecting architecture to solve the short path problem. Our experimental results show considerable performance gain can be achieved with reasonable area overhead.
Yu-Shih Su, Po-Hsien Chang, Shih-Chieh Chang 0001, TingTing Hwang
ACM Trans. Design Autom. Electr. Syst.1
2007 An Efficient Mechanism for Performance Optimization of Variable-Latency Designs
abstract
In many designs, the worst-case-delay path may never be exercised or may be exercised infrequently. For those designs, a strategy of optimizing a circuit for the worst-case conditions could lead to inefficient resource use. It is possible to improve the throughput of such circuits by introducing variable latency. One of the existing realizations of variable-latency design style is based on Telescopic Units. The design of the hold logic in telescopic units influences the circuit's throughput. In this paper, we show that the traditionally-designed hold logic in telescopic units may be inaccurate. We make use of the short path activation conditions to obtain more accurate hold logic than that commonly applied in the telescopic units. On average, our approach achieves a performance gain of 25.79% compared to 14.04%, which was reported in the previous works.
Yu-Shih Su, Da-Chung Wang, Shih-Chieh Chang 0001, Malgorzata Marek-Sadowska
DAC1