EDBT 2026 Demo / reviewers in the wild / expert
Chih-Chieh Hsiao
dblp:47/4658
· DBLP profile ↗
9ranked-venue papers
3as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-authorSystems, architecture and hardware · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
GPUs and heterogeneous computing · 50% Energy-efficient computing · 50% | |
| Computer graphics and multimedia
1 paper |
Rendering · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing › GPU scheduling
GPU thread scheduling |
0.2 | 1 | 2014 | An Adaptive Thread Scheduling Mechanism With Low-Power Register File for Mobile GPUs · IEEE Trans. Multim. 2014 |
GPUs and heterogeneous computing › embedded GPU
mobile GPU |
0.2 | 1 | 2014 | An Adaptive Thread Scheduling Mechanism With Low-Power Register File for Mobile GPUs · IEEE Trans. Multim. 2014 |
Energy-efficient computing
power management |
0.2 | 1 | 2014 | An Adaptive Thread Scheduling Mechanism With Low-Power Register File for Mobile GPUs · IEEE Trans. Multim. 2014 |
Energy-efficient computing › memory energy efficiency
register file energy reduction |
0.2 | 1 | 2014 | An Adaptive Thread Scheduling Mechanism With Low-Power Register File for Mobile GPUs · IEEE Trans. Multim. 2014 |
Rendering
real-time rendering |
0.1 | 1 | 2014 | An Adaptive Thread Scheduling Mechanism With Low-Power Register File for Mobile GPUs · IEEE Trans. Multim. 2014 |
Methods — techniques the papers use, named apart from their topics
multiple power modes · 0.4adaptive thread scheduling · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Demand look-ahead memory access scheduling for 3D graphics processing units
Chih-Chieh Hsiao, Min-Jen Lo, Slo-Li Chu |
Multim. Tools Appl. | 1 |
| 2014 | An Adaptive Thread Scheduling Mechanism With Low-Power Register File for Mobile GPUsabstractIn response to the remarkable increase in 3D applications in consumer electronics devices in recent years, graphics processing units (GPUs) have become widely available on mobile devices. These GPUs typically use hardware multithreaded shaders to improve their throughputs for real-time rendering, but they depend on duplicate register files to maintain the context of each hardware thread, increasing power consumption. However, the register usage of shading programs is often relatively low, which causes many registers to remain unused, thus wasting power. Long latency memory operations can also consume unnecessary power to activate registers. This study proposes a low-power register file with multiple power modes to reduce the power consumption of the register file. This study also presents an adaptive thread scheduling mechanism to achieve a tradeoff between the power consumption of the register file and frames per second (FPS). Results show that the average performance degradation from the proposed low-power register file is only 0.62%. The proposed adaptive thread scheduling has average under prediction ratio of 3.32%. The leakage reduction of the proposed low-power register file is 74.80%. This reduction can be improved to 81.49%, 82.22%, and 84.28% with adaptive thread scheduling at frame rates of 30, 25, and 20, respectively. Chih-Chieh Hsiao, Slo-Li Chu, Chiu-Cheng Hsieh |
IEEE Trans. Multim. | 1 |
| 2013 | Energy-aware hybrid precision selection framework for mobile GPUs
Chih-Chieh Hsiao, Slo-Li Chu, Chen-Yu Chen |
Comput. Graph. | 1 |
| 2011 | A Dual-Mode Unified Shader with Frame-Based Dynamic Precision Adjustment for Mobile GPUsabstractIn order to extend the life for battery driven mobile devices and maintain image quality, this paper presents a dual-mode unified shader for mobile GPUs, which consists of floating-point and fixed-point SIMD shader, for high quality or energy-saving rendering. Furthermore, in order to increase the image quality in fixed-point rendering, this paper proposes a frame-based dynamic precision adjustment scheme to select appropriate precision for different 3D scenes. The proposed design has following characteristics: I) high quality rendering with floating-point and fixed-point rendering for energy saving, II) a frame-based dynamic precision adjustment scheme to select appropriate precision for given scene, III) a workload-based scene change detection mechanism to re-select precision in time. Furthermore, this paper presents side by side comparison on performance, power and image quality between floating-point and fixed-point rendering in real world 3D games. The results of proposed shader in real world 3D games have 48.6% reduction in dynamic power and 33% faster in thread execution for a shader under energy saving mode in average. Furthermore, the rendered image qualities under proposed dynamic precision are insensitive to human eyes and the PSNR outperform related work for 2.37% in average. This reveals a way to use conventional fixed-point with dynamic precision to implement low power unified shader with quality rendering for such power limited devices. Slo-Li Chu, Chih-Chieh Hsiao, Chen-Yu Chen |
EUC | 2 |
| 2011 | An Energy-Efficient Unified Register File for Mobile GPUsabstractThe programmability of mobile GPUs have raised in recently years, where the shaders inside are instructed by shading programs for realistic 3D effects. The register files for a conventional high throughput multithreaded shader consumes 10% to 20% energy of it. However the register usages of shading program are quite low. In order to reduce the dynamic energy for register file in a multithreaded mobile shader, this paper proposed an unified register file design to reduce both dynamic and leakage energy of it. The result shows that proposed design reduces 85% of dynamic energy in a multithreaded register file. Furthermore, the proposed design reduces 59% of leakage energy and 25% of area with negligible performance degradation. Also, the energy savings in proposed designs are at least 75% more than related work. Slo-Li Chu, Chih-Chieh Hsiao, Chiu-Cheng Hsieh |
EUC | 2 |
| 2010 | OpenCL: Make Ubiquitous Supercomputing PossibleabstractDue to the dramatic requirements of 3D games and applications, graphics processing unit (GPU) or general-purpose graphics processing unit (GPGPU) have become required components in the modern computer systems. While these devices enable high parallelism with huge amount of processing elements, the utilization of their capabilities in general scientific applications are still low due to their difficult programming paradigms. Therefore an open standard, OpenCL, is proposed to provide universal APIs and programming paradigms for various GPUs and accelerators. In this study, it adopts several benchmarks, with various computation characteristics, to demonstrate the capabilities of OpenCL with several platforms. These programs are parallelized by OpenMP and OpenCL, and then targeted on several GPUs and conventional servers. This paper also provides an example to illustrate the migration of the given program, from OpenMP to OpenCL. The presented experimental results show that these inexpensive GPUs will lead better performance than servers if adopt OpenCL paradigms. It will be the preliminary milestone of cheap supercomputing by the acceleration of GPUs that can be obtained ubiquitously. Slo-Li Chu, Chih-Chieh Hsiao |
HPCC | 2 |
| 2009 | Design a Hardware Mechanism to Utilize Multiprocessors on a Uni-processor Operating System
Slo-Li Chu, Chih-Chieh Hsiao, Pin-Hua Chiu |
ICA3PP | 2 |
| 2008 | Memory Efficient Hierarchical Lookup Tables for Mass Arbitrary-Side Growing Huffman Trees DecodingabstractThis paper addresses the optimization problem of minimizing the number of memory access subject to a rate constraint for any Huffman decoding of various standard codecs. We propose a Lagrangian multiplier based penalty-resource metric to be the targeting cost function. To the best of our knowledge, there is few related discussion, in the literature, on providing a criterion to judge the approaches of entropy decoding under resource constraint. The existing approaches which dealt with the decoding of the single-side growing Huffman tree may not be memory-efficient for arbitrary-side growing Huffman trees adopted in current codecs. By grouping the common prefix part of a Huffman tree, in stead of the commonly used single-side growing Huffman tree, we provide a memory efficient hierarchical lookup table to speed up the Huffman decoding. Simulation results show that the proposed hierarchical table outperforms previous methods. A Viterbi-like algorithm is also proposed to efficiently find the optimal hierarchical table. More importantly, the Viterbi-like algorithm obtains the same results as that of the brute-force search algorithm. Sung-Wen Wang, Ja-Ling Wu, Shang-Chih Chuang, Chih-Chieh Hsiao, Yi-Shin Tung |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2006 | An Efficient Memory Construction Scheme for an Arbitrary Side Growing Huffman TableabstractBy grouping the common prefix of a Huffman tree, in stead of the commonly used single-side rowing Huffman tree (SGH-tree), we construct a memory efficient Huffman table on the basis of an arbitrary-side growing Huffman tree (AGH-tree) to speed up the Huffman decoding. Simulation results show that, in Huffman decoding, an AGH-tree based Huffman table is 2.35 times faster that of the Hashemian's method (an SGH-tree based one) and needs only one-fifth the corresponding memory size. In summary, a novel Huffman table construction scheme is proposed in this paper which provides better performance than existing construction schemes in both decoding speed and memory usage Sung-Wen Wang, Shang-Chih Chuang, Chih-Chieh Hsiao, Yi-Shin Tung, Ja-Ling Wu |
ICME | 3 |