Po-Kuan Huang

dblp:26/6007 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
0since 2021 · last 2007
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-authorSoftware engineering, systems software and programming languages · 3 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Electronic design automation · 76% Energy-efficient computing · 19% Embedded and real-time systems · 6%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization › compiler optimization
energy-aware compilation
0.112006
Leakage-aware intraprogram voltage scaling for embedded processors · DAC 2006
Electronic design automation › high-level synthesis
data path synthesis
0.112006
A Unified Theory of Timing Budget Management · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
Energy-efficient computing › voltage scaling
dynamic voltage scaling
0.112006
Leakage-aware intraprogram voltage scaling for embedded processors · DAC 2006
Electronic design automation
high-level synthesis
0.112006
A Unified Theory of Timing Budget Management · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
Electronic design automation › high-level synthesis
resource binding
0.112006
A Unified Theory of Timing Budget Management · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006
Electronic design automation
time budget assignment
0.112006
A Unified Theory of Timing Budget Management · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006

Methods — techniques the papers use, named apart from their topics

adaptive body biasing · 0.1min-cost flow · 0.1combinatorial optimization · 0.1
YearPublicationVenuePosition
2007 Efficient and scalable compiler-directed energy optimization for realtime applications
abstract
We present a compilation technique that targets realtime applications running on embedded processors with combined dynamic voltage scaling (DVS) and adaptive body biasing (ABB) capabilities. Considering the delay and energy penalty of switching between operating modes of the processor, our compiler judiciously inserts mode switch instructions in selected locations of the code and generates executable binary that is guaranteed to meet the deadline constraint. More importantly, our algorithm runs very fast and comes reasonably close to the theoretical limit of energy optimization using DVS+ABB. At 65 nm technology, we improve the energy dissipation of the generated code by an average of11.4% under deadline constraints. While our technique's improvement in energy dissipation over conventional DVS is marginal (3%) at 130nm, the average improvement continues to grow to 4.7%, 8.8% and 15.4% for 90nm, 65nm and 45nm technology nodes, respectively. Compared to a recent ILP-based competitor, we improve the runtime by more than three orders of magnitude, while producing improved results
Po-Kuan Huang, Soheil Ghiasi
DATE1
2007 Joint throughput and energy optimization for pipelined execution of embedded streaming applications
abstract
We present a methodology for synthesizing streaming applications, modeled as task graphs, for pipelined execution on multi-core architectures. We develop a task graph extraction and characterization framework that accurately determines the structure, computation and communication characteristics of application task graph from its specification in C. Furthermore, we develop a provably optimal algorithm that jointly balances the workload assigned to each core, and minimizes inter-core communication traffic. Experiment results show that our versatile method improves the through-put of streaming applications significantly under a variety of hardware configurations.
Po-Kuan Huang, Matin Hashemi, Soheil Ghiasi
LCTES1
2007 Efficient and scalable compiler-directed energy optimization for realtime applications
abstract
With continuing shrinkage of technology feature sizes, the share of leakage in total energy consumption of digital systems continues to grow. Coordinated supply voltage and body bias throttling enables the compiler to better optimize the total energy consumption of the system in future technology nodes. We present a compilation technique that targets realtime applications running on embedded processors with combined dynamic voltage scaling (DVS) and adaptive body biasing (ABB) capabilities. Considering the delay and energy penalty of switching between operating modes of the processor, our compiler judiciously inserts mode-switch instructions in selected locations of the code and generates executable binary that is guaranteed to meet the deadline constraint. More importantly, our algorithm runs very fast and comes reasonably close to the theoretical limit of energy optimization using DVS+ABB. At 65nm technology, we improve the energy dissipation of the generated code by an average of 33.20% under deadline constraints. While our technique's improvement in energy dissipation over conventional DVS is marginal (6.91%) at 130nm, the average improvement continues to grow to 13.19%, 22.97%, and 33.21% for 90nm, 65nm, and 45nm technology nodes, respectively. Compared to a recent ILP-based competitor, we improve the runtime by more than three orders of magnitude, while producing improved results.
Po-Kuan Huang, Soheil Ghiasi
ACM Trans. Design Autom. Electr. Syst.1
2006 Leakage-aware intraprogram voltage scaling for embedded processors
abstract
With scaling of technology feature sizes, the share of leakage in total power consumption of digital systems continues to grow. Conventional dynamic voltage scaling (DVS) techniques fail to accurately address theimpact of scaling on system power consumption and hence, are incapable ofachieving energy efficient solutions. To overcome this problem, we utilizeadaptive body biasing (ABB) to adjust transistors' threshold voltage at runtime. We develop a leakage-aware compilation methodology that targets embedded processors with both DVS and ABBcapabilities. Our technique has the unique advantage of jointly optimizing active and leakage energy dissipation. Considering the delay and energy penalty of switching between operating modes of the processor and under deadline constraint, our compiler improves the energy consumption of the generated code by average of 13.07% and up to 30.26% at 90nm. While our technique's improvement in energy dissipation over conventional DVS is marginal (4.54%) at 130nm,the average improvement continues to grow to 7.8%, 15.94% and 29.56% for 90nm, 65nm and 45 technology nodes, respectively.
Po-Kuan Huang, Soheil Ghiasi
DAC1
2006 Power-aware compilation for embedded processors with dynamic voltage scaling and adaptive body biasing capabilities
abstract
Traditionally, active power has been the primary source of power dissipation in CMOS designs. Although, leakage power is becoming increasingly more important as technology feature sizes continue to shrink, traditioinal power optimization techniques often neglect its contribution to total system power. In this paper, we present a power-aware compilation methodology that targets an embedded processor with both dynamic voltage scaling (DVS) and adaptive body biasing (ABB) capabilities. Our technique has the unique advantage of optimizing design power by jointly optimizing dynamic and leakage power dissipation. Considering the delay and energy penalty of swithching between processor modes, our compiler generates code with minimum power consumption under deadline constraints. Compared to not performing any optimization, or using DVS alone, our technique improves the power consumption of a number of embedded application kernels by 2 6 %, and 1 4 %, respectively.
Po-Kuan Huang, Soheil Ghiasi
DATE1
2006 A Unified Theory of Timing Budget Management
abstract
This paper presents a theoretical framework that solves optimally and in polynomial time many open problems in time budgeting. The approach unifies a large class of existing time-management paradigms. Examples include time budgeting for maximizing total weighted delay relaxation, minimizing the maximum relaxation, and min-skew time budget distribution. The authors develop a combinatorial framework through which we prove that many of the time-management problems can be transformed into a min-cost flow problem instance. The methodology is applied to intellectual-property-based datapath synthesis targeting field-programmable gate arrays. The synthesis flow maps the input operations to parameterized library modules during which different time budgeting policies have been applied. The techniques always improve the area requirement of the implemented test benches and consistently outperform a widely used competitor. The experiments verify that combining fairness and maximization objectives improves the results further as compared with pure maximum budgeting. The combined fairness and maximization objective improves the area by 25.8% and 28.7% in slice and LUT counts, respectively.
Soheil Ghiasi, Elaheh Bozorgzadeh, Po-Kuan Huang, Roozbeh Jafari, Majid Sarrafzadeh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2006 Probabilistic delay budget assignment for synthesis of soft real-time applications
abstract
Unlike their hard real-time counterparts, soft real-time applications are only expected to guarantee their "expected delay" over input data space. This paradigm shift calls for customized statistical design techniques to replace the conventional pessimistic worst case analysis methodologies. We present a novel statistical time-budgeting algorithm to translate the application expected delay constraint into its components' local delay constraints. We utilize the mathematical properties of the problem to quickly calculate the system expected delay and incrementally estimate the component utility variation with its timing relaxation. Our algorithm determines the optimal maximum weighted timing relaxation of an application under expected delay constraint. Experimental results on core-based synthesis of several multimedia applications targeting field-programmable gate arrays show that our technique always improves the design area. Furthermore, it consistently outperforms optimal time budgeting under hard real-time constraint, which is the best existing competitor. Design area improvements were up to 26% and averaged about 17% on several MediaBench applications.
Soheil Ghiasi, Po-Kuan Huang, Roozbeh Jafari
IEEE Trans. Very Large Scale Integr. Syst.2