Chiao-Ling Lung

dblp:36/8189 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Electronic design automation · 29% Hardware reliability and fault tolerance · 29% Energy-efficient computing · 29%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware reliability and fault tolerance
fault-tolerant design
0.322013
Through-Silicon Via Fault-Tolerant Clock Networks for 3-D ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2013
Fault-tolerant 3D clock network · DAC 2011
Electronic design automation
physical design
0.322013
Through-Silicon Via Fault-Tolerant Clock Networks for 3-D ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2013
Fault-tolerant 3D clock network · DAC 2011
Hardware reliability and fault tolerance › fault-tolerant design
TSV fault tolerance
0.322013
Through-Silicon Via Fault-Tolerant Clock Networks for 3-D ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2013
Fault-tolerant 3D clock network · DAC 2011
Embedded and real-time systems
real-time scheduling
0.212014
Thermal-Aware On-Line Scheduler for 3-D Many-Core Processor Throughput Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Energy-efficient computing › thermal management
thermal-aware scheduling
0.212014
Thermal-Aware On-Line Scheduler for 3-D Many-Core Processor Throughput Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Energy-efficient computing › thermal management
thermal-aware task allocation
0.212014
Thermal-Aware On-Line Scheduler for 3-D Many-Core Processor Throughput Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Energy-efficient computing
thermal management
0.212014
Thermal-Aware On-Line Scheduler for 3-D Many-Core Processor Throughput Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Electronic design automation › physical design
clock network synthesis
0.212013
Through-Silicon Via Fault-Tolerant Clock Networks for 3-D ICs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2013
Electronic design automation › physical design › clock network synthesis
clock tree synthesis
0.112011
Fault-tolerant 3D clock network · DAC 2011
Parallel and multicore computing › task scheduling
many-core scheduling
0.112014
Thermal-Aware On-Line Scheduler for 3-D Many-Core Processor Throughput Optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014

Methods — techniques the papers use, named apart from their topics

optimization modeling · 0.2online task migration · 0.2redundant clock trees · 0.2TSV fault-tolerant unit · 0.2
YearPublicationVenuePosition
2015 Soft-Error-Tolerant Design Methodology for Balancing Performance, Power, and Reliability
abstract
Soft error has become an important reliability issue in advanced technologies. To tolerate soft errors, solutions suggested in previous works incur significant performance and power penalties, especially when a design with fault-tolerant structures is overprotected. In this paper, we present a soft-error-tolerant design methodology to tradeoff performance, power, and reliability for different applications. First, four novel detection and correction flip-flop (FF) structures are proposed to provide different levels of tolerance capability against soft errors. Second, architecture-level vulnerability and logic-level susceptibility analyses are employed to identify weak FFs that can easily cause program execution errors. Third, an optimization framework is developed to synthesize the proposed four novel FF structures into weak and highly observable storage bits with the flexibility of trading off performance, power, and reliability. A five-stage pipeline RISC core (UniRISC) is adopted to demonstrate the usefulness of our methodology. Experimental results show that the proposed method can accomplish design goals by balancing performance, power, and reliability. For example, we can not only satisfy the reliability requirement that no more than five errors occur per one billion hours in a design but also reduce up to 87% performance overhead and 91% power overhead when compared with previous works.
Hsuan-Ming Chou, Ming-Yi Hsiao, Yi-Chiao Chen, Keng-Hao Yang, Jean Tsao, Chiao-Ling Lung, Shih-Chieh Chang 0001, Wen-Ben Jone, Tien-Fu Chen
IEEE Trans. Very Large Scale Integr. Syst.6
2014 Thermal-Aware On-Line Scheduler for 3-D Many-Core Processor Throughput Optimization
abstract
3-D many-core processor (3-D MCP) has become an emerging technology to tackle the power wall problem due to rapidly increasing number of transistors. However, when maximizing the throughput of 3-D MCP, which is expressed as a weighted sum of the speeds, due to the inherent heat removal limitation, thermal issues must be taken into consideration. Since the temperature of a core strongly depends on its location in the 3-D IC, a proper task allocation can alleviate the thermal problem and improve the throughput. Nevertheless, conventional techniques require computationally intensive thermal simulation, which prohibits its usage from the online application. In this paper, we propose an efficient online task allocation and task migration algorithm attempting to maximize the throughput of 3-D MCP simultaneously, considering unfinished tasks left from the last scheduling interval and new incoming tasks of this scheduling interval. The results of our experiments show that our proposed method achieves a 20.82X runtime speedup. These results are comparable to the exhaustive solutions obtained from optimization-modeling software LINGO. In addition, on average, our throughput results, with and without consideration of unfinished tasks, are only 4.39% and 0.69% worse, respectively, than that of the exhaustive method. In 128 task-to-core allocations, our method takes only 0.951 ms, which is 59.39 times faster than that of the previous work.
Cody Hao Yu, Chiao-Ling Lung, Yi-Lun Ho, Ruei-Siang Hsu, Ding-Ming Kwai, Shih-Chieh Chang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2013 Through-Silicon Via Fault-Tolerant Clock Networks for 3-D ICs
abstract
Clock network synthesis is one of the most important and challenging problems in 3-D ICs. The clock signals have to be delivered by through-silicon vias (TSVs) to different tiers with minimum skew. While there are a few related works in literature, none consider the reliability of TSVs in a clock tree. Accordingly, the failure of any TSV in the clock tree yields a bad chip. The naive solution using double-TSV can alleviate the problem, but the significant area overhead renders it less practical for large designs. In this paper, we propose a novel TSV fault-tolerant unit (TFU) to provide tolerance against TSV failures. The TFU makes use of the existing 2-D redundant trees designed for prebond testing, and thus has minimum area overhead. In addition, the number of TSVs in a TFU is also adjustable to allow flexibility during clock network synthesis. Compared with the conventional double TSV technique, the 3-D clock network constructed by TFUs can achieve 58% area overhead reduction with similar yield rate on an industrial case. To the best of the authors' knowledge, this is the first work in the literature that considers the fault tolerance of a 3-D clock network. It can be easily integrated with any bottom-up clock network synthesis algorithm.
Chiao-Ling Lung, Yu-Shih Su, Hsih-Hsiu Huang, Yiyu Shi 0001, Shih-Chieh Chang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2011 Fault-tolerant 3D clock network
abstract
Clock tree synthesis is one of the most important and challenging problems in 3D ICs. The clock signals have to be delivered by through-silicon vias (TSVs) to different tiers with minimum skew and latency. While there are a few related works in literature, none of them considers the reliability of TSVs. Accordingly, the failure of any TSV in the clock tree yields a bad chip. The naive solution using double-TSV can alleviate the problem. But the significant area overhead renders it less practical for large designs. In this paper, we propose a novel TSV fault-tolerant unit (TFU) that can provide tolerance against TSV failures in a 3D clock network. It makes use of the existing 2D redundant trees designed for pre-bond testing, and thus has minimum area overhead. Compared to the double TSV technique, the 3D clock network constructed by our TFUs can achieve 61% area reduction with 3.9% yield rate improvement on an industrial case. To the best of the authors' knowledge, this is the first practical work in literature that considers the fault tolerance of a 3D clock network.
Chiao-Ling Lung, Yu-Shih Su, Shih-Hsiu Huang, Yiyu Shi 0001, Shih-Chieh Chang 0001
DAC1
2011 Thermal-aware on-line task allocation for 3D multi-core processor throughput optimization
abstract
Three-dimensional integrated circuit (3D IC) has become an emerging technology in view of its advantages in packing density and flexibility in heterogeneous integration. The multi-core processor (MCP), which is able to deliver equivalent performance with less power consumption, is a candidate for 3D implementation. However, when maximizing the throughput of 3D MCP, due to the inherent heat removal limitation, thermal issues must be taken into consideration. Furthermore, since the temperature of a core strongly depends on its location in the 3D MCP, a proper task allocation helps to alleviate any potential thermal problem and improve the throughput. In this paper, we present a thermal-aware on-line task allocation algorithm for 3D MCPs. The results of our experiments show that our proposed method achieves 16.32X runtime speedup, and 23.18% throughput improvement. These are comparable to the exhaustive solutions obtained from optimization modeling software LINGO. On average, our throughput is only 0.85% worse than that of the exhaustive method. In 128 task-to-core allocations, our method takes only 0.932 ms, which is 57.74 times faster than the previous work.
Chiao-Ling Lung, Yi-Lun Ho, Ding-Ming Kwai, Shih-Chieh Chang 0001
DATE1
2010 Clock skew optimization considering complicated power modes
abstract
To conserve energy, a design which utilizes different power modes has been widely adopted. However, when a design has many different power modes, clock tree optimization (CTO) becomes very difficult. In this paper, we propose a two-level power-mode-aware CTO methodology. Among all different power modes, the chip-level CTO globally reduces clock skew among modules, whereas the module-level CTO reduces clock skew within a single module. Our experimental results show that the power-mode-aware CTO can achieve significant improvement in the worst-case condition with only a minor penalty in area.
Chiao-Ling Lung, Zi-Yi Zeng, Chung-Han Chou, Shih-Chieh Chang 0001
DATE1