Ja Chun Ku

dblp:14/3772 · DBLP profile ↗
← Back
13ranked-venue papers
10as first author
0since 2021 · last 2010
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 10 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Integrated circuit design · 41% Energy-efficient computing · 26% Memory systems · 16%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Integrated circuit design
digital circuit design
0.222008
Area Optimization for Leakage Reduction and Thermal Stability in Nanometer-Scale Technologies · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
On the Scaling of Temperature-Dependent Effects · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Integrated circuit design
low-power circuit design
0.222008
Area Optimization for Leakage Reduction and Thermal Stability in Nanometer-Scale Technologies · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
On the Scaling of Temperature-Dependent Effects · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Energy-efficient computing
leakage power reduction
0.122008
Area Optimization for Leakage Reduction and Thermal Stability in Nanometer-Scale Technologies · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
Thermal Management of On-Chip Caches Through Power Density Minimization · MICRO 2005
Electronic design automation › design optimization
area optimization
0.112008
Area Optimization for Leakage Reduction and Thermal Stability in Nanometer-Scale Technologies · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
Memory systems
cache design
0.112007
Variable latency caches for nanoscale processor · SC 2007
Processor architecture and microarchitecture
pipelining
0.112007
Variable latency caches for nanoscale processor · SC 2007
Integrated circuit design
technology scaling
0.112007
On the Scaling of Temperature-Dependent Effects · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Energy-efficient computing › thermal management
thermal-aware design
0.112007
On the Scaling of Temperature-Dependent Effects · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Memory systems › cache
on-chip cache
0.112005
Thermal Management of On-Chip Caches Through Power Density Minimization · MICRO 2005
Energy-efficient computing
thermal management
0.112005
Thermal Management of On-Chip Caches Through Power Density Minimization · MICRO 2005
Memory systems
cache
0.012007
Variable latency caches for nanoscale processor · SC 2007
Energy-efficient computing
leakage power
0.012007
On the Scaling of Temperature-Dependent Effects · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007

Methods — techniques the papers use, named apart from their topics

area as design parameter · 0.1simulation · 0.1alpha-power law · 0.1SPICE modeling · 0.1BSIM3 · 0.1row-level power gating · 0.1block permutation · 0.1
YearPublicationVenuePosition
2010 SACTA: A Self-Adjusting Clock Tree Architecture for Adapting to Thermal-Induced Delay Variation
abstract
Aggressive technology scaling down and low-power design techniques lead to uneven distributed power density, which translates into heat flow in the chips, causing significant temperature variations in both spatial and temporal terms. In order to mitigate the negative impacts of temperature variations on circuit timing, we propose SACTA, a self-adjusting clock tree architecture, which performs temperature-dependent dynamic clock skew scheduling to prevent timing violations in a pipelined circuit. The dynamic and adaptive features of SACTA are enabled by our proposed automatic temperature-adjustable skew buffers and temperature-insensitive skew buffers. These special delay elements are carefully tuned to ensure resilience of the entire circuit against temperature variation. To determine their configurations, we proposed an efficient and general clock tree design and optimization framework. Furthermore, we show that SACTA is applicable across a wide spectrum of circuits, including multi-${V}_{\rm dd}/{V}_{\rm th}$designs. Experimental results show that a pipeline supported by SACTA is able to prevent thermal-induced timing violations within a significantly larger range of operating temperatures (on average, the violation-free range can be enhanced by over 15$^{\circ}\hbox {C}$).
Jieyi Long, Ja Chun Ku, Seda Ogrenci Memik, Yehea I. Ismail
IEEE Trans. Very Large Scale Integr. Syst.2
2008 Area Optimization for Leakage Reduction and Thermal Stability in Nanometer-Scale Technologies
abstract
Traditionally, the minimum possible area of a very large scale integration (VLSI) layout is considered to be the best for delay and power minimization due to decreased interconnect capacitance. This paper, however, shows that the use of minimum area does not result in minimum power and/or delay in nanometer-scale technologies due to thermal effects and, in some cases, may cause thermal runaway. A methodology using area as a design parameter to reduce the leakage power and prevent thermal runaway is presented. A 16-bit adder example in 70-nm technology shows total power savings of 17% with 15% increase in area and no increase in delay. The power savings using this technique are expected to increase in future technologies.
Ja Chun Ku, Yehea I. Ismail
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2007 A self-adjusting clock tree architecture to cope with temperature variations
abstract
Ensuring resilience against environmental variations is becoming one of the great challenges of chip design. In this paper, we propose a self adjusting clock tree architecture, SACTA, to improve chip performance and reliability in the presence of on-chip temperature variations. SACTA performs temperature dependent dynamic clock skew scheduling to prevent timing violations in a pipelined circuit. We present an automatic temperature adjustable skew buffer design, which enables the adaptive feature of SACTA. Furthermore, we propose an efficient and general optimization framework to determine the configuration of these special delay elements. Experimental results show that a pipeline supported by SACTA is able to prevent thermal induced timing violations within a significantly larger range of operating temperatures (enhancing the violation-free range by as much as 45°C).
Jieyi Long, Ja Chun Ku, Seda Ogrenci Memik, Yehea I. Ismail
ICCAD2
2007 Attaining Thermal Integrity in Nanometer Chips
abstract
As technology moves into the nanometer era, undesirable trends such as increasing power density, leakage power, and temperature variation within a chip have made thermal effects emerge as a major bottleneck for further technology scaling. Thermal effects are no longer just considered as a reliability issue, but it has also become a fundamentally important and comprehensive problem that includes timing and power issues as well. This paper first overviews the impact of thermal effects on power and performance. Two thermal-aware design techniques, area optimization and low-power cache design, are briefly described. The paper states that there is still plenty of room for further improvement in the area of thermal-aware design.
Ja Chun Ku, Yehea I. Ismail
ISCAS1
2007 A Compact and Accurate Temperature-Dependent Model for CMOS Circuit Delay
abstract
With ever increasing power density and temperature variations within chips, it is very important to correctly model temperature effects on the devices in a compact way. In this paper, it is first shown that the temperature dependencies of the mobility and the saturation velocity need to be treated separately in modeling the current with temperature effects. Then, a new compact temperature-dependent model is presented for the on-current and transient behavior of a CMOS inverter based on the alpha-power law. The proposed model is shown to have an excellent agreement with BSIM3.
Ja Chun Ku, Yehea I. Ismail
ISCAS1
2007 Thermal-aware methodology for repeater insertion in low-power VLSI circuits
abstract
In this paper, the impact of thermal effects on low-power repeater insertion methodology is studied. An analytical methodology for thermal-aware repeater insertion that includes the electrothermal coupling between power, delay, and temperature is presented, and simulation results with global interconnect repeaters are discussed for 90nm and 65nm technology. Simulation results show that the proposed thermal-aware methodology can save 17.5% more power consumed by the repeaters compared to a thermal-unaware methodology for a given allowed delay penalty. In addition, the proposed methodology also results in a lower chip temperature, and thus, extra leakage power savings from other logic blocks.
Ja Chun Ku, Yehea I. Ismail
ISLPED1
2007 Variable latency caches for nanoscale processor
abstract
Variability is one of the important issues in nanoscale processors. Due to increasing importance of interconnect structures in submicron technologies, the physical location and phenomena such as coupling have an increasing impact on the latency of operations. Therefore, traditional view of rigid access latencies to components wil result in suboptimal architectures. In this paper, we devise a cache architecture with variable access latency. Particularly, we a) develop a non-uniform access level 1 data-cache, b) study the impact of coupling and physical location on level 1 data cache access latencies, and c) develop and study an architecture where the variable latency cache can be accessed while the rest of the pipeline remains synchronous. To find the access latency with different input address transitions and environmental conditions, we first build a SPICE model at a 45nm technology for a cache similar to that of the level 1 data cache of the Intel Prescott architecture. Motivated by the large difference between the worst and best case latencies and the shape of the distribution curve, we change the cache architecture to allow variable latency accesses. Since the latency of the cache is not known at the time of instruction scheduling, we also modify the functional units with the addition of special queues that will temporarily store the dependent instructions and allow the data to be forwarded from the cache to the functional units correctly. Simulations based on SPEC2000 benchmarks show that our variable access latency cache structure can reduce the execution time by as much as 19.4% and 10.7% on average compared to a conventional cache architecture.
Serkan Ozdemir, Arindam Mallik, Ja Chun Ku, Gokhan Memik, Yehea I. Ismail
SC3
2007 On the Scaling of Temperature-Dependent Effects
abstract
With ever increasing power density and temperature variations within chips, it is very important to correctly model temperature effects on the devices in a compact way and to predict their scaling. In this paper, it is first shown that the temperature dependences of the mobility and the saturation velocity need to be treated separately in modeling the current with the temperature effects. A new compact temperature- dependent model for the ON-current is presented based on the alpha-power law and is verified with BSIM3. Then, the scaling of the ON-current temperature dependence is discussed. It is also shown in this paper that the temperature effects will have an increasing impact on repeater-insertion methodology. Furthermore, the temperature-dependence scaling of the leakage current is analyzed. It is shown that its temperature dependence decreases with technology scaling, but temperature-aware power- reduction techniques will actually save larger fraction of the total power due to the increasing dominance of leakage power.
Ja Chun Ku, Yehea I. Ismail
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2007 Thermal-Aware Methodology for Repeater Insertion in Low-Power VLSI Circuits
abstract
In this paper, the impact of thermal effects on low-power repeater insertion methodology is studied. An analytical methodology for thermal-aware repeater insertion that includes the electrothermal coupling between power, delay, and temperature is presented, and simulation results with global interconnect repeaters are discussed for 90- and 65-nm technology. Simulation results show that the proposed thermal-aware methodology can save 17.5% more power consumed by the repeaters compared to a thermal-unaware methodology for a given allowed delay penalty. In addition, the proposed methodology also results in a lower chip temperature, and thus, extra leakage power savings from other logic blocks.
Ja Chun Ku, Yehea I. Ismail
IEEE Trans. Very Large Scale Integr. Syst.1
2007 Thermal Management of On-Chip Caches Through Power Density Minimization
abstract
Various architectural power reduction techniques have been proposed for on-chip caches in the last decade. In this paper, we first show that these power reduction techniques can be suboptimal when thermal effects are considered. Then, we propose a thermal-aware cache power-down technique that minimizes the power density of the active parts by turning off alternating rows of memory cells instead of entire banks. The decrease in the power density lowers the temperature, which then exponentially reduces the leakage. Thus, leakage power of the active parts is reduced in addition to the power eliminated from the parts that are turned off. Simulations based on SPEC2000, NetBench, and MediaBench applications in a 70-nm technology show that the proposed thermal-aware architecture can reduce the total energy consumption by 53% compared to a conventional cache, and 14% compared to a cache architecture with thermal-unaware power reduction scheme. Second, we show a block permutation scheme that can be used during the design of the caches to maximize the distance between blocks with consecutive addresses. Because of spatial locality, blocks with consecutive addresses are likely to be accessed within a short time interval. By maximizing the distance between such blocks, we minimize the power density of the hot spots in the cache, and hence reduce the peak temperature. This, in return, results in an average leakage power reduction of 8.7% compared to a conventional cache without affecting the dynamic power and the latency. Overall, both of our architectures add no extra run-time penalty compared to the thermal-unaware power reduction schemes, yet they result in a significant reduction in the total energy consumption of a cache
Ja Chun Ku, Serkan Ozdemir, Gokhan Memik, Yehea I. Ismail
IEEE Trans. Very Large Scale Integr. Syst.1
2006 Area optimization for leakage reduction and thermal stability in nanometer scale technologies
abstract
Traditionally, minimum possible area of a VLSI layout is considered the best for delay and power minimization due to decreased interconnect capacitance. This paper shows however that the use of minimum area does not result in the minimum power and/or delay in nanometer scale technologies due to thermal effects, and in some cases, may result in thermal runaway. A methodology using area as a design parameter to reduce the leakage power, and prevent thermal runaway is presented. A 16-bit adder example in a 70nm technology shows a total power savings of 17% with 15% increase in area, and no increase in delay. The power savings using this technique are expected to increase in future technologies.
Ja Chun Ku, Yehea I. Ismail
ASP-DAC1
2006 Power density minimization for highly-associative caches in embedded processors
abstract
Caches are essential components in embedded processors, taking up a significant fraction of the chip area and power. As a result of the relatively large size and infrequent activity, leakage power of caches is becoming an important problem. There exist a number of power density minimization schemes that distribute the activity evenly among computational entities, thereby lowering the temperature to reduce the leakage power. In this paper, we first present various power density minimization schemes for highly-associative caches in embedded processors via access distribution. It is then suggested that they should be used in conjunction with other power-down techniques to be more effective. We show that conventional power-down techniques for on-chip caches can be suboptimal if thermal effects are ignored, and propose a thermal-aware power-down technique that minimizes power density of the active parts. Simulations based on MediaBench, NetBench, and MiBench applications in a 70nm technology show that the proposed thermal-aware schemes can improve leakage power savings of a conventional power-down technique by 8.5% on average, and up to 23%.
Ja Chun Ku, Serkan Ozdemir, Gokhan Memik, Yehea I. Ismail
ACM Great Lakes Symposium on VLSI1
2005 Thermal Management of On-Chip Caches Through Power Density Minimization
abstract
Various architectural power reduction techniques have been proposed for on-chip caches in the last decade. However, these techniques mostly ignore the effects of temperature on the power consumption. In this paper, first we show that these power reduction techniques can be suboptimal when thermal effects are considered. Particularly, we propose a thermal-aware cache power-down technique that minimizes the power density of the active parts by turning off alternating rows of memory cells instead of entire banks. The decrease in the power density lowers the temperature, which in return, reduces the leakage of the active parts. Simulations based on SPEC2000 benchmarks in a 70nm technology show that the proposed thermal-aware architecture can reduce the total energy consumption by 53% compared to a conventional cache, and 14% compared to a cache architecture with thermal-unaware power reduction scheme. Second, we show a block permutation scheme that can be used during the design of caches to maximize the distance between blocks with consecutive addresses. By maximizing the distance between consecutively accessed blocks, we minimize the power density of the hot spots in the cache, and hence reduce the peak temperature. This, in return, results in an average leakage power reduction of 8.7% compared to a conventional cache without affecting the dynamic power and the latency. Overall, both of our architectures add no extra run-time penalty compared to the thermal-unaware power reduction schemes, yet they reduce the total energy consumption of a conventional cache by 53% and 5.6% on average, respectively.
Ja Chun Ku, Serkan Ozdemir, Gokhan Memik, Yehea I. Ismail
MICRO1