VLDB 2026 Research / reviewers in the wild / expert
Jenq Kuen Lee
dblp:42/3701 · also Jenq-Kuen Lee
· DBLP profile ↗
69ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0001-9919-6258ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 55 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 7Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Work-in-Progress: Extending a RISC-V Core with Sub-FP8 Support for Machine LearningabstractModern deep learning workloads demand massive computational resources, making energy-efficient computing paradigms essential. Currently, compiler and hardware support for emerging formats like FP8 (E5M2/E4M3), FP6 (E3M2/E2M3), and FP4 remain limited in the RISC-V ecosystem. This paper presents a scalar sub-FP8 ISA extension for RISC-V, with full LLVM toolchain and hardware support. Our design decouples format specification from instructions using the fcsr to store format information. Post-synthesis evaluation confirms hardware efficiency, using only an increase of 0.9% registers and 2.3% additional LUTs in CVA6, while delivering 20× lower cache miss rates for FP8 GEMM versus FP32 operations. In addition, we introduce future work on a sub-FP8 vector extension, with a focus on support for FP8, FP6, and FP4. Kathryn Chapman, Fu-Jian Shen, Jhih-Kuan Lin, Jenq Kuen Lee |
CODES+ISSS | 4 |
| 2025 | Support of MISRA C++ Analyzer for Reliability of Embedded SystemsabstractCyber-Physical Systems (CPS) are increasingly used in many complex applications, such as autonomous delivery drones, the automotive CPS design, power grid control systems, and medical robotics. However, existing programming languages lack certain design patterns for CPS designs, including temporal semantics and concurrency models. Future research directions may involve programming language extensions to support CPS designs. However, JSF++, MISRA, and MISRA C++ are providing specifications intended to increase the reliability of safety-critical systems. This article also describes the development of rule checkers based on the MISRA C++ specification using the Clang open-source tool, which allows for the annotation of code and the easy extension of the MISRA C++ specification to other programming languages and systems. This is potentially useful for future CPS language research extensions to work with reliability software specifications using the Clang tool. Experiments were performed using key C++ benchmarks to validate our method in comparison with the well-known Coverity commercial tool. We illustrate key rules related to class, inheritance, template, overloading, and exception handling. Open-source benchmarks that violate the rules detected by our checkers are also illustrated. A random graph generator is further used to generate diamond case with multiple inheritance test data for our software validations. The experimental results demonstrate that our method can provide information that is more detailed than that obtained using Coverity for nine open-source C++ benchmarks. Since the Clang tool is widely used, it will further allow developers to annotate their own extensions. Che-Chia Lin, Wei-Hsu Chu, Chia-Hsuan Chang, Hui-Hsin Liao, Chun-Chieh Yang, Jenq Kuen Lee, Yi-Ping You, Tien-Yuan Hsieh |
ACM Trans. Cyber Phys. Syst. | 6 |
| 2025 | Optimizing computer vision algorithms with TVM on VLIW architecture based on RVV
Meng-Shiun Yu, Hao-Chun Chang, Chong-Teng Wang, Yu-Wei Tien, Tai-Liang Chen, Jenq Kuen Lee |
J. Supercomput. | 6 |
| 2024 | Low DRAM Memory Access and Flexible Dataflow Convolutional Neural Network Accelerator based on RISC-V Custom InstructionabstractDeep convolutional neural networks has been widely used in several applications. However, the huge computational complexity and data access times hinder its application in edge devices. Previous works target to design a specific fixed dataflow. However, several researches point out that there are no dataflow that can be optimal across all layers or models. In this paper, first we propose an flexible dataflow accelerator which can reconfigure to weight stationary or output stationary dataflow at every layer to increase hardware utilization and data reuse. Besides, we design RISC-V custom instructions to encode the dataflow configurations. Last, we proposed a on-the-fly pooling method to compute the max pooling layer right after the convolutional layer to reduce off-chip memory access. By reconfiguring the dataflow, we improve 3.75x and 1.18x DRAM access amounts in VGG16 compared with [1], [2] respectively. Besides, we maintain a high utilization rate of 99.12%. The proposed accelerator can not only reconfigure the dataflow of each layer but also achieve high throughput, high area efficiency, and high power efficiency. The accelerator implemented in the 40nm process reaching 256 GOPS throughput with 1000 MHz, 136.9 GOPS/mm2area efficiency with 1.87 mm2area, 1.014 TOPS/W power efficiency with 252.51 mW power. Yu-Jen Chang, Ching-Te Chiu, Ming-Long Huang, Geng-Ming Liang, Chao-Lin Lee, Jenq Kuen Lee, Ping-Yu Hsieh, Wei-Chih Lai |
ISCAS | 7 |
| 2023 | Accelerating AI performance with the incorporation of TVM and MediaTek NeuroPilotabstractThe continuing prominence of machine learning has led to an increased focus on enhancing the inference performance of edge devices to reduce latency and improve efficiency.Two widely adopted strategies for accelerating computational performance are quantisation and the utilisation of AI hardware accelerators.Each type of accelerator or inference engine offers distinct advantages, with accelerators primarily designed to optimise neural network operations.In this paper, we present an innovative method for integrating TVM's quantisation flow with the MediaTek Neuropilot AI accelerator.We outline the process of converting the TVM relay intermediate-representation quantised neural network dialect model to a tensor-oriented quantisation format, with the aim of harnessing the full potential of both TVM and MediaTek NeuroPilot.This integration enables more efficient neural network inference while preserving the accuracy of the results.We assessed the effectiveness of our proposed integration by conducting a series of experiments and comparing the performance of our approach with that of TVM equipped with an autotuning mechanism.The findings indicate that our approach substantially outperforms TVM in both floatingpoint model inference and quantised model inference, with inference speedups of up to 11× and up to 70×, respectively.These results underscore the potential of our approach in accelerating AI performance across a diverse range of applications and edge devices.Moreover, a key contribution of our work is providing a valuable practical method for other hardware companies interested in integrating TVM with their own accelerators to achieve performance gains. Chao-Lin Lee, Chun-Ping Chung, Sheng-Yuan Cheng, Jenq Kuen Lee, Robert Lai |
Connect. Sci. | 4 |
| 2023 | Auto-tuning Fixed-point Precision with TVM on RISC-V Packed SIMD ExtensionabstractToday, as deep learning (DL) is applied more often in daily life, dedicated processors such as CPUs and GPUs have become very important for accelerating model executions. With the growth of technology, people are becoming accustomed to using edge devices, such as mobile phones, smart watches, and VR devices in their daily lives. A variety of technologies using DL are gradually being applied to these edge devices. However, there is a large number of computations in DL. It faces a challenging problem how to provide solutions in the edge devices. In this article, the proposed method enables a flow with the RISC-V Packed extension (P extension) in TVM. TVM, an open deep learning compiler for neural network models, is growing as a key infrastructure for DL computing. RISC-V is an open instruction set architecture (ISA) with customized and flexible features. The Packed-SIMD extension is a RISC-V extension that enables subword single-instruction multiple-data (SIMD) computations in RISC-V architectures to support fallback engines in AI computing. In the proposed flow, a fixed-point type that is supported by an integer of 16-bit type and saturation instructions is added to replace the original 32-bit float type. In addition, an auto-tuning method is proposed to use a uniform selector mechanism (USM) to find the binary point position for fixed-point type use. The tensorization feature of TVM can be used to optimize specific hardware such as subword SIMD instructions with RISC-V P extension. With our experiment on the Spike simulator, the proposed method with the USM can improve performance by approximately 2.54 to 6.15× in terms of instruction counts with little accuracy loss. Chun-Chieh Yang, Hui-Hsin Liao, Yuan-Ming Chang, Jenq Kuen Lee |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2022 | Efficient Realization of Decision Trees for Real-Time InferenceabstractFor timing-sensitive edge applications, the demand for efficient lightweight machine learning solutions has increased recently. Tree ensembles are among the state-of-the-art in many machine learning applications. While single decision trees are comparably small, an ensemble of trees can have a significant memory footprint leading to cache locality issues, which are crucial to performance in terms of execution time. In this work, we analyze memory-locality issues of the two most common realizations of decision trees, i.e., native and if-else trees. We highlight that both realizations demand a more careful memory layout to improve caching behavior and maximize performance. We adopt a probabilistic model of decision tree inference to find the best memory layout for each tree at the application layer. Further, we present an efficient heuristic to take architecture-dependent information into account thereby optimizing the given ensemble for a target computer architecture. Our code-generation framework, which is freely available on an open-source repository, produces optimized code sessions while preserving the structure and accuracy of the trees. With several real-world data sets, we evaluate the elapsed time of various tree realizations on server hardware as well as embedded systems for Intel and ARM processors. Our optimized memory layout achieves a reduction in execution time up to 75 % execution for server-class systems, and up to 70 % for embedded systems, respectively. Kuan-Hsun Chen, Chiahui Su, Christian Hakert, Sebastian Buschjäger, Chao-Lin Lee, Jenq Kuen Lee, Katharina Morik, Jian-Jia Chen |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2021 | Support NNEF execution model for NNAPI
Yuan-Ming Chang, Chia-Yu Sung, Yu-Chien Sheu, Meng-Shiun Yu, Min-Yih Hsu, Jenq Kuen Lee |
J. Supercomput. | 6 |
| 2021 | NNBlocks: a Blockly framework for AI computing
Tai-Liang Chen, Meng-Shiun Yu, Jenq Kuen Lee |
J. Supercomput. | 4 |
| 2020 | Experiment and enabled flow for GPGPU-Sim simulators with fixed-point instructions
Chao-Lin Lee, Min-Yih Hsu, Bing-Sung Lu, Ming-Yu Hung, Jenq Kuen Lee |
J. Syst. Archit. | 5 |
| 2018 | Architecture and Compiler Support for GPUs Using Energy-Efficient Affine Register FilesabstractA modern GPU can simultaneously process thousands of hardware threads. These threads are grouped into fixed-size SIMD batches executing the same instruction on vectors of data in a lockstep to achieve high throughput and performance. The register files are huge due to each SIMD group accessing a dedicated set of vector registers for fast context switching, and consequently the power consumption of register files has become an important issue. One proposed solution is to replace some of the vector registers by scalar registers, as different threads in a same SIMD group operate on scalar values and so the redundant computations and accesses of these scalar values can be eliminated. However, it has been observed that a significant number of registers containing affine vectors υ such that υ[ i ] = b + i × s can be represented by base b and stride s . Therefore, this article proposes an affine register file design for GPUs that is energy efficient due to it reducing the redundant executions of both the uniform and affine vectors. This design uses a pair of registers to store the base and stride of each affine vector and provides specific affine ALUs to execute affine instructions. A method of compiler analysis has been developed to detect scalars and affine vectors and annotate instructions for facilitating their corresponding scalar and affine computations. Furthermore, a priority-based register allocation scheme has been implemented to assign scalars and affine vectors to appropriate scalar and affine register files. Experimental results show that this design was able to dispatch 43.56% of the computations to scalar and affine ALUs when using eight scalar and four affine registers per warp. This resulted in the current design also reducing the energy consumption of the register files and ALUs to 21.86% and 26.54%, respectively, and it reduced the overall energy consumption of the GPU by an average of 5.18%. Shao-Chung Wang, Li-Chen Kan, Chao-Lin Lee, Yuan-Shin Hwang, Jenq Kuen Lee |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2017 | Analyzing OpenCL 2.0 workloads using a heterogeneous CPU-GPU simulatorabstractHeterogeneous CPU-GPU systems have recently emerged as an energy-efficient computing platform. A robust integrated CPU-GPU simulator is essential to facilitate researches in this direction. While few integrated CPU-GPU simulators are available, similar tools that support OpenCL 2.0, a widely used new standard with promising heterogeneous computing features, are currently missing. In this paper, we extend the existing integrated CPU-GPU simulator, gem5-gpu, to support OpenCL 2.0. In addition, we conduct experiments on the extended simulator to see the impact of new features introduced by OpenCL 2.0. Our OpenCL 2.0 compatible simulator is successfully validated against a state-of-the-art commercial product, and is expected to help boost future studies in heterogeneous CPU-GPU systems. Ren-Wei Tsai, Shao-Chung Wang, Kun-Chih Chen, Po-Han Wang 0001, Hsiang-Yun Cheng, Yi-Chung Lee, Sheng-Jie Shu, Chun-Chieh Yang, Min-Yih Hsu, Li-Chen Kan, Chao-Lin Lee, Tzu-Chieh Yu, Rih-Ding Peng, Chia-Lin Yang, Yuan-Shin Hwang, Jenq Kuen Lee, Shiao-Li Tsao, Ouhyoung Ming |
ISPASS | 17 |
| 2017 | Enabling PoCL-based runtime frameworks on the HSA for OpenCL 2.0 support
Yuan-Ming Chang, Shao-Chung Wang, Chun-Chieh Yang, Yuan-Shin Hwang, Jenq Kuen Lee |
J. Syst. Archit. | 5 |
| 2016 | Vector data flow analysis for SIMD optimizations on OpenCL programsabstractSummary Multi‐core systems equipped with micro processing units and accelerators such as digital signal processors (DSPs) and graphics processing units (GPUs) have become a major trend in processor design in recent years in attempts to meet ever‐increasing application performance requirements. Open Computing Language (OpenCL) is one of the programming languages that include new extensions proposed to exploit the computing power of these kinds of processors. Among the newly extended language features, the single‐instruction multiple‐data (SIMD) linguistics and vector types are added to OpenCL to exploit hardware features of the accelerators. The addition makes it necessary to consider how traditional compiler data flow analysis can be adopted to meet the optimization requirements of vector linguistics. In this paper, we propose a calculus framework to support the data flow analysis of vector constructs for OpenCL programs that compilers can use to perform SIMD optimizations. We model OpenCL vector operations as data access functions in the style of mathematical functions. We then show that the data flow analysis for OpenCL vector linguistics can be performed based on the data access functions. Based on the information gathered from data flow analysis, we illustrate a set of SIMD optimizations on OpenCL programs. The experimental results incorporating our calculus and our proposed compiler optimizations show that the proposed SIMD optimizations can provide average performance improvements of 22% on x86 CPUs and 4% on advanced micro devices GPUs. For the selected 15 benchmarks, 11 of them are improved on x86 CPUs, and six of them are improved on advanced micro devices GPUs. The proposed framework has the potential to be used to construct other SIMD optimizations on OpenCL programs. Copyright © 2015 John Wiley & Sons, Ltd. Yu-Te Lin, Jenq Kuen Lee |
Concurr. Comput. Pract. Exp. | 2 |
| 2016 | Translating the ARM Neon and VFP instructions in a binary translatorabstractSummary Binary translation attempts to emulate one instruction set with another on the same or different platforms. The important technique is widely used in modern software. Vector and floating‐point instructions are widely used in many applications, including multimedia, graphics, and gaming. Although these instructions are usually simulated with software in a binary translator, it is important to support them such that the host single‐instruction, multiple‐data (SIMD) and floating‐point hardware are efficiently used during emulation. We report our design and implementation of the emulation of ARM Neon and vector floating point (VFP) instructions in the machine‐code‐to‐low‐level‐virtual‐machine (MC2LLVM) binary translator. The Neon and VFP instructions are first translated into carefully chosen sequences of LLVM intermediate representation (IR), and later, the IR sequences are optimized and translated into the host native binary by the existing LLVM backend. Because MC2LLVM makes use of the vector and floating‐point types in LLVM IR, the generated host native binary can take full advantage of the vector and floating‐point functional units, if present, of the host machine. To be fully compliant with Neon and VFP instruction sets, all the features are supported, including the flush‐to‐zero mode, default not a number mode, and floating‐point exceptions. The experimental results show that code generated by MC2LLVM with the Neon and VFP extensions achieves an average speedup of 1.174× in SPEC 2006 benchmark suites and exhibits a floating‐point throughput of 12.05× in LINPACK, compared with code generated by MC2LLVM without the Neon and VFP extensions. Furthermore, MC2LLVM is 3.36× faster than QEMU for processing Neon/VFP instructions. Copyright © 2016 John Wiley & Sons, Ltd. Yu-Chuan Guo, Wuu Yang, Jiunn-Yeu Chen, Jenq Kuen Lee |
Softw. Pract. Exp. | 4 |
| 2015 | The Design and Experiments of A SID-Based Power-Aware Simulator for Embedded Multicore SystemsabstractEmbedded multicore systems are playing increasingly important roles in the design of consumer electronics. The objective of such systems is to optimize both performance and power characteristics of mobile devices. However, currently there are no power metrics supporting popular application design platforms (such as SID) that application developers use to develop their applications. This hinders the ability of application developers to optimize power consumption. In this article we present the design and experiments of a SID-based power-aware simulation framework for embedded multicore systems. The proposed power estimation flow includes two phases: IP-level power modeling and power-aware system simulation. The first phase employs PowerMixer IP to construct the power model for the processor IP and other major IPs, while the second phase involves a power abstract interpretation method for summarizing the simulation trace, then, with a CPE module, estimating the power consumption based on the summarized trace information and the input of IP power models. In addition, a Manager component is devised to map each digital signal processor (DSP) component to a host thread and maintain the access to shared resources. The aim is to maintain the simulation performance as the number of simulated DSP components increases. A power-profiling API is also supported that developers of embedded software can use to tune the granularity of power-profiling for a specific code section of the target application. We demonstrate via case studies and experiments how application developers can use our SID-based power simulator for optimizing the power consumption of their applications. We characterize the power consumption of DSP applications with the DSPstone benchmark and discuss how compiler optimization levels with SIMD intrinsics influence the performance and power consumption. A histogram application and an augmented-reality application based on human-face-based RMS (recognition, mining, and synthesis) application are deployed as running examples on multicore systems to demonstrate how our power simulator can be used by developers in the optimization process to illustrate different views of power dissipations of applications. Cheng-Yen Lin, Chung-Wen Huang, Chi-Bang Kuan, Shi-Yu Huang, Jenq Kuen Lee |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2014 | The design of LLVM-based shader compiler for embedded architectureabstractThe increasing hand-held devices has resulted in widespread use of Apps. To sustain the Apps performance, most developers use Graphic Processing Units (GPUs) as hardware accelerator recently. As a result, GPUs become the basic system requirement in the smart hand-held devices. With the improvement of GPUs techniques, the rendering pipeline has become programmable by using the shading language proposed from DirectX and OpenGL. A shader compiler is essential to compile the shading language into GPU assembly in order to run on programmable rendering pipelines. In this thesis, we describe a case study for the process to implement an OpenGL-ES 2.0 shader compiler for new GPU designed by an ITRI team. The compiler is based on LLVM and OpenGL framework Mesa. We use the front-end from Mesa that translates the OpenGL-ES shading language (GLSL ES) 2.0 to TGSI intermediate form and the Gallium3D device driver framework in Mesa will translate the TGSI to LLVM bitcode. Finally, we use the LLVM as backend to compile the bitcode to assembly NV gpu program4. In our preliminary experimental results, we use the gpu simualtor provided from ITRI and GLES benchmark to show that our LLVM based shader compiler is able to generate reliable codes for basic OpenGL-ES programs. Li-Wei Kuo, Chun-Chieh Yang, Jenq Kuen Lee, Shau-Yin Tseng |
ICPADS | 3 |
| 2014 | On the and-or-scheduling problemsabstractIn the and-or scheduling model, a project consists of several tasks. Each task has a duration attribute. A task can be performed only when all of its requirements are satisfied. After a task is completed, more requirements become satisfied. A characteristic of the AOscheduling projects is that a requirement may be satisfied in several ways. Several questions concerning AOscheduling might be interesting, including whether the project can be completed, the earliest time a project can be completed, the minimal number of processors needed to complete the project, and assigning tasks to processors, etc. We use Petri nets and segment graphs to analyze AOscheduling projects. Wuu Yang, Ming-Hsiang Huang, Jenq Kuen Lee |
ICPADS | 3 |
| 2014 | Register spilling via transformed interference equations for PAC DSP architectureabstractSUMMARY Digital signal processors (DSPs) with very long instruction word (VLIW) data‐path architectures are increasingly being deployed on embedded devices for multimedia processing applications. To reduce the power consumption and design cost of VLIW DSP processors, distributed register files and multibank register architectures are being adopted to reduce the number of read and write ports associated with register files, which presents new challenges for devising compiler optimization schemes. This paper addresses the issues of reducing the spill code for a VLIW DSP with distributed register files. Spill code produced by register allocation is traditionally handled by memory spills, but the multibank register‐file architecture provides the opportunity to spill‐out register values onto different register banks. We present a conceptual framework based on the universal and the proxy interference graphs to model the live ranges of registers for spilling codes to different register banks. Heuristic algorithms are then developed on the basis of this concept. By heuristically estimating the register pressure for each register file, we treat different register banks as optional spilling locations in addition to traditional spilling to memory. Experiments were performed on the parallel architecture core VLIW DSP with distributed register files by incorporating our proposed optimization schemes into an Open64‐based compiler. The experimental results show that our approach can improve the performances on average for DSPStone and MiBench benchmarks with spilling cases by 7.1% and 21.6%, respectively, compared with the one always handling spill code in memory. Copyright © 2013 John Wiley & Sons, Ltd. Chung-Ju Wu, Chia-Han Lu, Jenq Kuen Lee |
Concurr. Comput. Pract. Exp. | 3 |
| 2014 | Achieving spilling-friendly register file assignment for highly distributed register files
Chia-Han Lu, Wen-Li Shih, Chung-Ju Wu, Jenq Kuen Lee |
J. Supercomput. | 4 |
| 2014 | Compiler Optimization for Reducing Leakage Power in Multithread BSP ProgramsabstractMultithread programming is widely adopted in novel embedded system applications due to its high performance and flexibility. This article addresses compiler optimization for reducing the power consumption of multithread programs. A traditional compiler employs energy management techniques that analyze component usage in control-flow graphs with a focus on single-thread programs. In this environment the leakage power can be controlled by inserting on and off instructions based on component usage information generated by flow equations. However, these methods cannot be directly extended to a multithread environment due to concurrent execution issues. This article presents a multithread power-gating framework composed of multithread power-gating analysis (MTPGA) and predicated power-gating (PPG) energy management mechanisms for reducing the leakage power when executing multithread programs on simultaneous multithreading (SMT) machines. Our multithread programming model is based on hierarchical bulk-synchronous parallel (BSP) models. Based on a multithread component analysis with dataflow equations, our MTPGA framework estimates the energy usage of multithread programs and inserts PPG operations as power controls for energy management. We performed experiments by incorporating our power optimization framework into SUIF compiler tools and by simulating the energy consumption with a post-estimated SMT simulator based on Wattch toolkits. The experimental results show that the total energy consumption of a system with PPG support and our power optimization method is reduced by an average of 10.09% for BSP programs relative to a system without a power-gating mechanism on leakage contribution set to 30%; and the total energy consumption is reduced by an average of 4.27% on leakage contribution set to 10%. The results demonstrate our mechanisms are effective in reducing the leakage energy of BSP multithread programs. Wen-Li Shih, Yi-Ping You, Chung-Wen Huang, Jenq Kuen Lee |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2013 | Compilers for Low Power with Design Patterns on Embedded Multicore SystemsabstractMinimization of power dissipation can be considered at algorithmic, compilers, architectural, logic, and circuit levels. Recent research trends for multicore programming models have come to the direction that parallel design patterns can be a solution to develop multicore applications. As parallel design patterns are with regularity, we view this as a great opportunity to exploit power optimizations in the software layer. In this paper, we present case studies to investigate compilers for low power with parallel design patterns on embedded multicore systems. We evaluate two major parallel design patterns, Pipe and Filter and MapReduce with Iterator. Our work, attempts to devise power optimization schemes in compilers by exploiting the opportunities of the recurring patterns of embedded multicore programs. In all two cases of the patterns investigated, the common recurring patterns of programs are exploited to seek the opportunity for compiler optimizations for low power. Proposed optimization schemes are rate-based optimization for Pipe and Filter pattern and early-exit power optimization for MapReduce with Iterator pattern. Our experiment is based on a power simulator simulating a heterogeneous multicore system under SID simulation framework. In our experiments, a finite impulse response (FIR) program with Pipe and Filter pattern and an image recognition application applied MapReduce with Iterator pattern are evaluated by incorporating our proposed power optimization schemes for each pattern. Significant power reductions are observed in all two cases. With the case study, we present a direction for power optimizations that one can further identify additional key design patterns for embedded multicore systems to explore power optimization opportunities via compilers. Cheng-Yen Lin, Chi-Bang Kuan, Jenq Kuen Lee |
ICPP | 3 |
| 2012 | Compiler supports for VLIW DSP processors with SIMD intrinsicsabstractSUMMARY To sustain growing multimedia workload, modern digital signal processing (DSP) processors are commonly equipped with subword instructions to accelerate signal processing. Besides subword, functional units of very long instruction word (VLIW) DSP processors can also be employed to process multiple data streams in parallel. However, because of power and area concerns, many embedded VLIW DSP processors adopt distributed register files to reduce read/write ports and wire connection by privatizing register files for clusters and even for functional units. The distributed design presents great challenges to compilers in distributing single instruction, multiple data (SIMD) workload to functional units. In this paper, we address the issue in supporting SIMD parallelism on VLIW DSP processors with subword instructions and distributed register files. Currently, industrial practices have adopted intrinsics that enable developers to utilize hardware resources and compete with hand‐coded assembly in performance. However, it is still an open issue to provide such a solution for VLIW DSP processors with distributed register files. In this work, we provide SIMD intrinsics to allow programmers to write highly optimized codes by following given programming guides. In addition, an enhanced register allocation scheme and data replication optimizations are devised to enable efficient code generation. In our experiments, DSPstone benchmark and a set of H.264 kernels are used to evaluate the proposed programming and optimization schemes. The result shows that by combining SIMD intrinsics and compiler optimizations, one is able to obtain remarkable performance improvements, speedups of 2.9 and 3.5 for DSPstone and H.264 kernels, respectively. Copyright © 2011 John Wiley & Sons, Ltd. Chi-Bang Kuan, Jenq Kuen Lee |
Concurr. Comput. Pract. Exp. | 2 |
| 2012 | Parallelization of Belief Propagation on Cell Processors for Stereo VisionabstractMarkov random field models provide a robust formulation for the stereo vision problem of inferring three-dimensional scene geometry from two images taken from different viewpoints. One of the most advanced algorithms for solving the associated energy minimization problem in the formulation is belief propagation (BP). Although BP provides very accurate results in solving stereo vision problems, the high computational cost of the algorithm hinders it from real-time applications. In recent years, multicore architectures have been widely adopted in various industrial application domains. The high computing power of multicore processors provides new opportunities to implement stereo vision algorithms. This article examines and extracts the parallelisms in the BP method for stereo vision on multicore processors. This article shows that parallelism of the algorithm can be efficiently utilized on multicore processors. The results show that parallelization on multicore processors provides a speedup for the BP algorithm of almost 15 times compared to the single-processor implementation on the PPE of the Cell BE. The experimental results also indicate that a frame rate of 6.5 frames/second is possible when implementing the parallelized BP algorithm on the multicore processor of Cell BE with one PPE and six SPEs. Kun-Yuan Hsieh, Chi-Hua Lai, Shang-Hong Lai, Jenq Kuen Lee |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2012 | Case study: stereo vision experiments with multi-core software API on embedded MPSoC environments
Jia-Jhe Li, Chung-Kai Chen, Tung-Yu Wu, Jenq Kuen Lee |
J. Supercomput. | 4 |
| 2012 | Instruction scheduling methods and phase ordering framework for VLIW DSP processors with distributed register files
Chung-Ju Wu, Yu-Te Lin, Jenq Kuen Lee |
J. Supercomput. | 3 |
| 2012 | Support of Probabilistic Pointer Analysis in the SSA FormabstractProbabilistic pointer analysis (PPA) is a compile-time analysis method that estimates the probability that a points-to relationship will hold at a particular program point. The results are useful for optimizing and parallelizing compilers, which need to quantitatively assess the profitability of transformations when performing aggressive optimizations and parallelization. This paper presents a PPA technique using the static single assignment (SSA) form. When computing the probabilistic points-to relationships of a specific pointer, a pointer relation graph (PRG) is first built to represent all of the possible points-to relationships of the pointer. The PRG is transformed by a sequence of reduction operations into a compact graph, from which the probabilistic points-to relationships of the pointer can be determined. In addition, PPA is further extended to interprocedural cases by considering function related statements. We have implemented our proposed scheme including static and profiling versions in the Open64 compiler, and performed experiments to obtain the accuracy and scalability. The static version estimates branch probabilities by assuming that every conditional is equally likely to be true or false, and that every loop executes 10 times before terminating. The profiling version measures branch probabilities dynamically from past program executions using a default workload provided with the benchmark. The average errors for selected benchmarks were 3.80 percent in the profiling version and 9.13 percent in the static version. Finally, SPEC CPU2006 is used to evaluate the scalability, and the result indicates that our scheme is sufficiently efficient in practical use. The average analysis time was 35.59 seconds for an average of 98,696 lines of code. Ming-Yu Hung, Peng-Sheng Chen, Yuan-Shin Hwang, Roy Dz-Ching Ju, Jenq Kuen Lee |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2011 | Enable OpenCL Compiler with Open64 InfrastructuresabstractAs microprocessors evolve into heterogeneous architectures with multi-cores of MPUs and GPUs, programming model supports become important for programming such architectures. To address this issue, OpenCL is proposed. Currently, most of OpenCL implementations take LLVM as their infrastructures. This presents an opportunity to demonstrate whether OpenCL can be effectively implemented on other compiler infrastructures. For example, Open64, which is another open source compiler and known to generate efficient codes for microprocessors, can contribute further to performance improvements and enhancing the adoption of heterogeneous computing based on OpenCL. In this paper, we describe the flow to enable an OpenCL compiler based on Open64 infrastructures for ATI GPUs. Our work includes the extension of the front-end parser for OpenCL, the generation of high-level intermediate representations with OpenCL linguistics, performing high-level optimization, and finally applying OpenCL specific optimization for code generations. Preliminary experimental results show that our compiler based on Open64 is able to generate efficient codes for OpenCL programs. Yu-Te Lin, Shao-Chung Wang, Wen-Li Shih, Kun-Yuan Hsieh, Jenq Kuen Lee |
HPCC | 5 |
| 2009 | pTest: An adaptive testing tool for concurrent software on embedded multicore processorsabstractMore and more processor manufacturers have launched embedded multicore processors for consumer electronics products because such processors provide high performance and low power consumption to meet the requirements of mobile computing and multimedia applications. To effectively utilize computing power of multicore processors, software designers interest in using concurrent processing for such architecture. The master-slave model is one of the popular programming models for concurrent processing. Even if it is a simple model, the potential concurrency faults and unreliable slave systems still lead to anomalies of entire system. In this paper, we present an adaptive testing tool called pTest to stress test a slave system and to detect the synchronization anomalies of concurrent software in the master-slave systems on embedded multicore processors. We use a probabilistic finite-state automaton(PFA) to model the test patterns for stress testing and shows how a PFA can be applied to pTest in practice. Shou-Wei Chang, Kun-Yuan Hsieh, Jenq Kuen Lee |
DATE | 3 |
| 2009 | Efficient multiple virtual view generation based on reduced depth stereo image for advanced autostereoscopic displaysabstractRecent development of autostereoscopic displays demands the synthesis of more virtual views in wider baseline, an inevitable trend for future 3DTV systems. However, current standard file formats, e.g. multi-view coding (MVC) and image plus depth, encounter challenging problems to meet the request. Therefore, we propose a new file format, called reduced depth stereo image (RDSI), which saves the color and depth images of the left view and the disoccluded regions in the right view. Based on RDSI, rendering virtual images from parallel viewpoints along the baseline can be simplified as view interpolation that is not only very efficient but also consistent in disocclusion regions. Through experiments on both simulated and real data, we demonstrate the superior performance with several quantitative assessments for the adopted file format and the rendering algorithm. The results suggest RDSI as a better choice to meet the demands for the online synthesis of many virtual views in wide baseline for advanced autostereoscopic displays. Chia-Ming Cheng, Shu-Jyuan Lin, Shang-Hong Lai, Jenq Kuen Lee |
ICME | 4 |
| 2009 | LC-GRFA: global register file assignment with local consciousness for VLIW DSP processors with non-uniform register filesabstractAbstract Embedded processors developed within the past few years have employed novel hardware designs to reduce the ever‐growing complexity, power dissipation, and die area. Although using a distributed register file architecture is considered to have less read/write ports than using traditional unified register file structures, it presents challenges in compilation techniques to generate efficient codes for such architectures. This paper presents a novel scheme for register allocation that includes global and local components on a VLIW DSP processor with distributed register files whose port access is highly restricted. In the scheme, an optimization phase performed prior to conventional global/local register allocation, named global/local register file assignment (RFA), is used to minimize various register file communication costs. A heuristic algorithm is proposed for global RFA to make suitable decisions based on local RFA. Experiments were performed by incorporating our schemes on a novel VLIW DSP processor with non‐uniform register files. The results indicate that the compilation based on our proposed approach delivers significant performance improvements, compared with the solution without using our proposed global register allocation scheme. Copyright © 2008 John Wiley & Sons, Ltd. Chia-Han Lu, Yung-Chia Lin, Yi-Ping You, Jenq Kuen Lee |
Concurr. Comput. Pract. Exp. | 4 |
| 2008 | Enabling Streaming Remoting on Embedded Dual-Core ProcessorsabstractDual-core processors (and, to an extent, multicore processors) have been adopted in recent years to provide platforms that satisfy the performance requirements of popular multimedia applications. This architecture comprises groups of processing units connected by various interprocess communication mechanisms such as shared memory, memory mapping interrupts, mailboxes, and channel-based protocols. The associated challenges include how to provide programming models and environments for developing streaming applications for such platforms. In this paper, we present middleware called streaming RPC for supporting a streaming-function remoting mechanism on asymmetric dual-core architectures. This middleware has been implemented both on an experimental platform known asthe PAC dual-core platform and in TI OMAP dual-core environments. We also present an analytic model of streaming equations to optimize the internal handshaking for our proposed streaming RPC. The usage and efficiency of the proposed methodology are demonstrated in a JPEG decoder, MP3 decoder, and QCIF H.264 decoder. The experimental results show that our approach improves the performance of the decoders of JPEG, MP3, and H.264 by 24%, 38%, and 32% on PAC, respectively. The communication load of internal handshaking has also been reduced compared to the naive use of RPC over embedded dual-core systems. The experiments also show that the performance improvement can also be achieved on OMAP dual-core platforms. Kun-Yuan Hsieh, Yen-Chih Liu, Po-Wen Wu, Shou-Wei Chang, Jenq Kuen Lee |
ICPP | 5 |
| 2008 | Mobile Java RMI support over heterogeneous wireless networks: A case study
Chung-Kai Chen, Cheng-Wei Chen, Chien-Tan Ko, Jenq Kuen Lee, Jyh-Cheng Chen |
J. Parallel Distributed Comput. | 4 |
| 2008 | Software architecture design for streaming Java RMI
Chih-Chieh Yang, Chung-Kai Chen, Yu-Hao Chang, Kai-Hsin Chung, Jenq Kuen Lee |
Sci. Comput. Program. | 5 |
| 2007 | Enabling compiler flow for embedded VLIW DSP processors with distributed register filesabstractHigh-performance and low-power VLIW DSP processors are in-creasingly deployed on embedded devices to process video and multimedia applications. For reducing power and cost in designs of VLIW DSP processors, distributed register files and multi-bank register architectures are being adopted to eliminate the amount of read/write ports in register files. This presents new challenges for devising compiler optimization schemes for such architectures. In this paper, we address the compiler optimization issues for PAC ar-chitecture, which is a 5-way issue DSP processor with distributed register files. We present an integrated flow to address several phases of compiler optimizations in interacting with distributed register files and multi-bank register files in the layer of instruc-tion scheduling, software pipelining, and data flow optimizations. Our experiments on a novel 32-bit embedded VLIW DSP (known as the PAC DSP core) exhibit the state of the art performance for embedded VLIW DSP processors with distributed register files by incorporating our proposed schemes in compilers. Chung-Kai Chen, Ling-Hua Tseng, Shih-Chang Chen, Young-Jia Lin, Yi-Ping You, Chia-Han Lu, Jenq Kuen Lee |
LCTES | 7 |
| 2007 | PALF: compiler supports for irregular register files in clustered VLIW DSP processorsabstractAbstract A wide variety of register file architectures—developed for embedded processors—have recently been used with the aim of reducing power dissipation and die size, in contrast with the traditional unified register file structures. This article presents a novel register allocation scheme for a clustered VLIW DSP, which is designed with distinctively banked register files in which port access is highly restricted. Whilst the organization of the register files is designed to decrease power consumption by using fewer port connections, the cluster‐based design makes register access across clusters an additional issue, and the switched‐access nature of the register file demands further investigation into the use of optimizing register assignment as a means of increasing instruction‐level parallelism. We propose a heuristic algorithm, namedping‐pong aware local favorable(PALF) register allocation, to obtain a register allocation that is expected to better utilize irregular register file architectures. The results of experiments performed using a compiler based on the Open Research Compiler (ORC) showed significant performance improvement over the original ORC's approach, which is considered to be an optimized approach for common register file architectures. Copyright © 2007 John Wiley & Sons, Ltd. Yung-Chia Lin, Yi-Ping You, Jenq Kuen Lee |
Concurr. Comput. Pract. Exp. | 3 |
| 2007 | Switching supports for stateful object remoting on network processors
Chung-Kai Chen, Yu-Hao Chang, Yu-Tin Chen, Chih-Chieh Yang, Jenq Kuen Lee |
J. Supercomput. | 5 |
| 2007 | Energy-aware scheduling and simulation methodologies for parallel security processors with multiple voltage domains
Yung-Chia Lin, Yi-Ping You, Chung-Wen Huang, Jenq Kuen Lee, Wei-Kuan Shih, TingTing Hwang |
J. Supercomput. | 4 |
| 2007 | Compilation for compact power-gating controlsabstractPower leakage constitutes an increasing fraction of the total power consumption in modern semiconductor technologies due to the continuing size reductions and increasing speeds of transistors. Recent studies have attempted to reduce leakage power using integrated architecture and compiler power-gating mechanisms. This approach involves compilers inserting instructions into programs to shut down and wake up components, as appropriate. While early studies showed this approach to be effective, there are concerns about the large amount of power-control instructions being added to programs due to the increasing amount of components equipped with power-gating controls in SoC design platforms. In this article we present a sink-n-hoist framework for a compiler to generate balanced scheduling of power-gating instructions. Our solution attempts to merge several power-gating instructions into a single compound instruction, thereby reducing the amount of power-gating instructions issued. We performed experiments by incorporating our compiler analysis and scheduling policies into SUIF compiler tools and by simulating the energy consumption using Wattch toolkits. The experimental results demonstrate that our mechanisms are effective in reducing the amount of power-gating instructions while further reducing leakage power compared to previous methods. Yi-Ping You, Chung-Wen Huang, Jenq Kuen Lee |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2006 | Power Aware H.264/AVC Video Player on PAC Dual-Core SoC Platform
Jia-Ming Chen, Chih-Hao Chang, Shau-Yin Tseng, Jenq Kuen Lee, Wei-Kuan Shih |
EUC | 4 |
| 2006 | PAC DSP Core and Application ProcessorsabstractThis paper provides an overview of the parallel architecture core (PAC) project led by SoC Technology Center of Industrial Technology Research Institute (STC/ITRI) in Taiwan. The background of PAC project, a brief introduction to PAC core technologies, PAC SoC development suite, PAC benchmarks, and applications are presented. The main objective of the PAC development plan is to enhance industrial development competitiveness in the core technology related to key components, especially for portable multimedia applications David Chih-Wei Chang, I-Tao Liao, Jenq Kuen Lee, Shau-Yin Tseng, Chein-Wei Jen |
ICME | 3 |
| 2006 | Integrating Compiler and System Toolkit Flow for Embedded VLIW DSP ProcessorsabstractTo support high-performance and low-power for multimedia applications and for hand-held devices, embedded VLIW DSP processors are of research focus. With the tight resource constraints, distributed register files, variable-length encodings for instructions, and special data paths are frequently adopted. This creates challenges to deploy software toolkits for new embedded DSP processors. This article presents our methods and experiences to develop software and toolkit flows for PAC (parallel architecture core) VLIW DSP processors. Our toolkits include compilers, assemblers, debugger and DSP micro-kernels. We first retarget open research compiler (ORC) and toolkit chains for PAC VLIW DSP processor and address the issues to support distributed register files and ping-pong data paths for embedded VLIW DSP processors. Second, the linker and assembler are able to support variable length encoding schemes for DSP instructions. In addition, the debugger and DSP micro-kernel were designed to handle dual-core environments. The footprint of micro-kernel is also around 10K to address the code-size issues for embedded devices. We also present the experimental result in the compiler framework by incorporating software pipeline (SWP) policies for distributed register files in PAC architecture. Results indicated that our compiler framework gains performance improvement around 2.5 times against the code generated without our proposed optimizations Kun-Yuan Hsieh, Yung-Chia Lin, Chung-Ju Wu, Wen-Li Shih, Shih-Chang Chen, Chung-Kai Chen, Chien-Ching Huang, Yi-Ping You, Jenq Kuen Lee |
RTCSA | 10 |
| 2006 | Compilers for leakage power reductionabstractPower leakage constitutes an increasing fraction of the total power consumption in modern semiconductor technologies. Recent research efforts indicate that architectures, compilers, and software can be optimized so as to reduce the switching power (also known as dynamic power) in microprocessors. This has lead to interest in using architecture and compiler optimization to reduce leakage power (also known as static power) in microprocessors. In this article, we investigate compiler-analysis techniques that are related to reducing leakage power. The architecture model in our design is a system with an instruction set to support the control of power gating at the component level. Our compiler provides an analysis framework for utilizing instructions to reduce the leakage power. We present a framework for analyzing data flow for estimating the component activities at fixed points of programs whilst considering pipeline architectures. We also provide equations that can be used by the compiler to determine whether employing power-gating instructions in given program blocks will reduce the total energy requirements. As the duration of power gating on components when executing given program routines is related to the number and complexity of program branches, we propose a set of scheduling policies and evaluate their effectiveness. We performed experiments by incorporating our compiler analysis and scheduling policies into SUIF compiler tools and by simulating the energy consumptions on Wattch toolkits. The experimental results demonstrate that our mechanisms are effective in reducing leakage power in microprocessors. Yi-Ping You, Chingren Lee, Jenq Kuen Lee |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2005 | System-level design space exploration for security processor prototyping in analytical approachesabstractThe customization of architectures in designing the security processor-based systems typically involves timeconsuming simulation and sophisticated analysis in the exploration of design spaces. In this paper, we present an analytical modeling strategy for synoptically exploring of the candidate architectures of security processor-based systems. of We demonstrate examples to employ our analytical models for design space explorations of embedded security systems to deal with scalability issues and architecture constraints. The experiments with the cycle-accurate simulation exhibit the applicability of analytical modeling: average prediction error is less than 10% while speed improvement is in several orders of magnitude. Yung-Chia Lin, Chung-Wen Huang, Jenq Kuen Lee |
ASP-DAC | 3 |
| 2005 | A sink-n-hoist framework for leakage power reductionabstractPower leakage constitutes an increasing fraction of the total power consumption in modern semiconductor technologies. Recent research efforts have tried to integrate architecture and compiler solutions to employ power-gating mechanisms to reduce leakage power. This approach is to have compilers perform data-flow analysis and insert instructions at programs to shut down and wake up components whenever appropriate for power reductions. While this approach has been shown to be effective in early studies, there are concerns for the amount of power-control instructions being added to programs with the increasing amount of components equipped with power-gating control in a SoC design platform. In this paper, we present a Sink-N-Hoist framework in the compiler solution to generate balanced scheduling of power-gating instructions. Our solution will attempt to merge power-gating instructions as one compound instruction. Therefore, it will reduce the amount of power-gating instructions issued.We perform experiments by incorporating our compiler analysis and scheduling policies into SUIF compiler tools and by simulating the energy consumptions on Wattch toolkits. The experimental results demonstrate that our mechanisms are effective in reducing the amount of power-gating instructions while further in reducing leakage power compared to previous methods. Yi-Ping You, Chung-Wen Huang, Jenq Kuen Lee |
EMSOFT | 3 |
| 2005 | Efficient Switching Supports of Distributed .NET Remoting with Network ProcessorsabstractDistributed object-oriented environments have become important platforms for parallel and distributed service frameworks. Among distributed object-oriented software, .NET Remoting provides a language layer of abstractions for performing parallel and distributed computing in .NET environments. In this paper, we present our methodologies in supporting .NET Remoting over meta-clustered environments. We take the advantage of the programmability of network processors to develop the content-based switch for distributing workloads generated from remote invocations in .NET. Our scheduling mechanisms include stateful supports for .NET Remoting services. In addition, we also propose scheduling policy to incorporate workflow models as the models are now incorporated in many of tools of grid architectures. Experiments done at clusters with IXP 1200 network processors show that our scheme can significantly enhance the system throughput (up to 55%) compared to NLB method when the traffic is heavy. Our schemes are effective in supporting the switching of .NET Remoting computations over meta-cluster environments. Chung-Kai Chen, Yu-Hao Chang, Cheng-Wei Chen, Yu-Tin Chen, Chih-Chieh Yang, Jenq Kuen Lee |
ICPP | 6 |
| 2005 | Support and optimization of Java RMI over a Bluetooth environmentabstractAbstract Distributed object‐oriented platforms are increasingly important over wireless environments for providing frameworks for collaborative computations and for managing a large pool of distributed resources. Due to limited bandwidths and heterogeneous architectures of wireless devices, studies are needed into supporting object‐oriented frameworks over heterogeneous wireless environments and optimizing system performance. In our research work, we are working towards efficiently supporting object‐oriented environments over heterogeneous wireless environments. In this paper, we report the issues and our research results related to the efficient support of Java RMI over a Bluetooth environment. In our work, we first implement support for Java RMI over Bluetooth protocol stacks, by incorporating a set of protocol stack layers for Bluetooth developed by us (which we call JavaBT) and by supporting the L2CAP layer with sockets that support the RMI socket. In addition, we model the cost for the access patterns of Java RMI communications. This cost model is used to guide the formation and optimizations of the scatternets of a Java RMI Bluetooth environment. In our approach, we employ the well‐known BTCP algorithm to observe initial configurations for the number of piconets. Using the communication‐access cost as a criterion, we then employ a spectral‐bisection method to cluster the nodes in a piconet and then use a bipartite matching scheme to form the scatternet. Experimental results with the prototypes of Java RMI support over a Bluetooth environment show that our scatternet‐formation algorithm incorporating an access‐cost model can further optimize the performances of such as system. Copyright © 2005 John Wiley & Sons, Ltd. Pu-Chen Wei, Chung-Hsin Chen, Cheng-Wei Chen, Jenq Kuen Lee |
Concurr. Pract. Exp. | 4 |
| 2004 | Efficient support of java RMI over heterogeneous wireless networksabstractDistributed object-oriented platforms are increasingly important over wireless environments to provide frameworks for collaborative computations and for managing a large pool of distributed resources. For beyond 3G environments, distributed object-oriented platforms can provide the framework and toolkits for application developments with heterogeneous wireless environments. In this paper, we present our support for Java RMI over Bluetooth, GPRS, and WLAN environments. We propose a software mechanism which can be used to dynamically adapt RMI over different networks with optimization-related strategies. This is an important middleware for component communications. We will show how to employ Java Dynamic Proxy and exception handling techniques to help perform roaming and resource scheduling among heterogeneous wireless environments. Java Grande benchmarks are used to demonstrate that our RMI implementations over GPRS, WLAN, and Bluetooth environments are effective in supporting parallel and distributed control of Java layers over heterogeneous wireless environments. Cheng-Wei Chen, Chung-Kai Chen, Jyh-Cheng Chen, Chien-Tan Ko, Jenq Kuen Lee, Hong-Wei Lin, Wang-Jer Wu |
ICC | 5 |
| 2004 | Specification and Architecture Supports for Component Adaptations on Distributed EnvironmentsabstractSummary form only given. With the arrival of the new computing paradigm in addressing autonomous systems for heterogeneous distributed architectures, distributed component technologies face challenges ahead. One of the key issues is how one can have component models adapt and respond to environment changes autonomously. We argue that additional annotation specifications for components are needed to advance this process. We first present additional annotation specifications for components. This information can then be retrieved by Java introspection and represented in DAML+OIL language, which is based on the RDF schema and the XML syntax. Based on the specifications, we then present a component management service (CMS) model and architecture to address the specification and composition issues of components on heterogeneous distributed architectures. Experimental results show significant performance improvements with our support of component adaptations in all cases. Our work presents a major advance for areas related to specifications and compositions of distributed components. Chung-Kai Chen, Cheng-Wei Chen, Jenq Kuen Lee |
IPDPS | 3 |
| 2004 | Case study: an infrastructure for C/ATLAS environments with object-oriented design and XML representation
Cheng-Wei Chen, Jenq Kuen Lee |
J. Syst. Softw. | 2 |
| 2004 | Support and optimization for parallel sparse programs with array intrinsics of Fortran 90
Rong-Guey Chang, Tyng-Ruey Chuang, Jenq Kuen Lee |
Parallel Comput. | 3 |
| 2004 | Interprocedural Probabilistic Pointer AnalysisabstractWhen performing aggressive optimizations and parallelization to exploit features of advanced architectures, optimizing and parallelizing compilers need to quantitatively assess the profitability of any transformations in order to achieve high performance. Useful optimizations and parallelization can be performed if it is known that certain points-to relationships would hold with high or low probabilities. For instance, if the probabilities are low, a compiler could transform programs to perform data speculation or partition iterations into threads in speculative multithreading, or it would avoid conducting code specialization. Consequently, it is essential for compilers to incorporate pointer analysis techniques that can estimate the possibility for every points-to relationship that it would hold during the execution. However, conventional pointer analysis techniques do not provide such quantitative descriptions and, thus, hinder compilers from more aggressive optimizations, such as thread partitioning in speculative multithreading, data speculations, code specialization, etc. We address this issue by proposing a probabilistic points-to analysis technique to compute the probability of every points-to relationship at each program point. A context-sensitive interprocedural algorithm has been implemented based on the iterative data flow analysis framework, and has been incorporated into SUIF and MachSUIF. Experimental results show this technique can estimate the probabilities of points-to relationships in benchmark programs with reasonable small errors, about 4.6 percent on average. Furthermore, the current implementation cannot disambiguate heap and array elements. The errors are further significantly reduced when the future implementation incorporates techniques to disambiguate heap and array elements. Peng-Sheng Chen, Yuan-Shin Hwang, Roy Dz-Ching Ju, Jenq Kuen Lee |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2003 | Compiler support for speculative multithreading architecture with probabilistic points-to analysisabstractSpeculative multithreading (SpMT) architecture can exploit thread-level parallelism that cannot be identified statically. Speedup can be obtained by speculatively executing threads in parallel that are extracted from a sequential program. However, performance degradation might happen if the threads are highly dependent, since a recovery mechanism will be activated when a speculative thread executes incorrectly and such a recovery action usually incurs a very high penalty. Therefore, it is essential for SpMT to quantify the degree of dependences and to turn off speculation if the degree of dependences passes certain thresholds. This paper presents a technique that quantitatively computes dependences between loop iterations and such information can be used to determine if loop iterations can be executed in parallel by speculative threads. This technique can be broken into two steps. First probabilistic points-to analysis is performed to estimate the probabilities of points-to relationships in case there are pointer references in programs, and then the degree of dependences between loop iterations is computed quantitatively. Preliminary experimental results show compiler-directed thread-level speculation based on the information gathered by this technique can achieve significant performance improvement on SpMT. Peng-Sheng Chen, Ming-Yu Hung, Yuan-Shin Hwang, Roy Dz-Ching Ju, Jenq Kuen Lee |
PPoPP | 5 |
| 2003 | Segmented Alignment: An Enhanced Model to Align Data Parallel Programs of HPF
Gwan-Hwan Hwang, Cheng-Wei Chen, Jenq Kuen Lee, Roy Dz-Ching Ju |
J. Supercomput. | 3 |
| 2003 | Compiler optimization on VLIW instruction scheduling for low powerabstractIn this article, we investigate compiler transformation techniques regarding the problem of scheduling VLIW instructions aimed at reducing power consumption of VLIW architectures in the instruction bus. The problem can be categorized into two types: horizontal scheduling and vertical scheduling. For the case of horizontal scheduling, we propose a bipartite-matching scheme for instruction scheduling. We prove that our greedy bipartite-matching scheme always gives the optimal switching activities of the instruction bus for given VLIW instruction scheduling policies. For the case of vertical scheduling, we prove that the problem is NP-hard, and we further propose a heuristic algorithm to solve the problem. Our experiment is performed on Alpha-based VLIW architectures and an ATOM simulator, and the compiler incorporated in our proposed schemes is implemented based on SUIF and MachSUIF. Experimental results of horizontal scheduling optimization show an average 13.30% reduction with four-way issue architecture and an average 20.15% reduction with eight-way issue architecture for transitional activities of the instruction bus as compared with conventional list scheduling for an extensive set of benchmarks. The additional reduction for transitional activities of the instruction bus from horizontal to vertical scheduling with window size four is around 4.57 to 10.42%, and the average is 7.66%. Similarly, the additional reduction with window size eight is from 6.99 to 15.25%, and the average is 10.55%. Chingren Lee, Jenq Kuen Lee, TingTing Hwang, Shi-Chun Tsai |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2001 | Probabilistic Inference Schemes for Sparsity Structures of Fortran 90 Array IntrinsicsabstractIn this paper, we address the issues of partitioning sparse arrays whose non-zero elements are distributed non-uniformly. We consider inference schemes for Fortran 90 array intrinsics so that the non-zero structure of the output array can be deduced from the non-zero structures of the input arrays. Experiments are conducted to measure the effectiveness of our method with the Harwell-Boeing sparse matrix collection. We also demonstrate that, given the sparsity structures of the source arrays and with the help of our inference schemes, one can predict the performance differences among a collection of equivalent Fortran 90 code for sample on-line analytical processing (OLAP). The experiments are performed on an IBM SP2 cluster with the library support of our sparse array intrinsics. Rong-Guey Chang, Jia-Shin Li, Jenq Kuen Lee, Tyng-Ruey Chuang |
ICPP | 3 |
| 2001 | Array Operation Synthesis to Optimize HPF Programs on Distributed Memory Machines
Gwan-Hwan Hwang, Jenq Kuen Lee, Roy Dz-Ching Ju |
J. Parallel Distributed Comput. | 2 |
| 2001 | Parallel Sparse Supports for Array Intrinsic Functions of Fortran 90
Rong-Guey Chang, Tyng-Ruey Chuang, Jenq Kuen Lee |
J. Supercomput. | 3 |
| 1999 | Compiler Optimizations for Parallel Sparse Programs with Array Intrinsics of Fortran 90abstractIn our recent work, we have been working on providing parallel sparse supports for array intrinsics of Fortran 90. Our supporting library uses a two-level design. In the low-level routines, it requires the input sparse matrices to be specified with compression/distribution schemes for array functions. In the high-level representations, sparse array functions are overloaded with Fortran 90 array intrinsic interfaces so that programmers need not be concerned about low-level details. This raises a very interesting optimization problem in the strategies to transform high-level representations to low-level routines by automatic selections and supplies of distribution and compression schemes for sparse arrays. We propose solutions to this optimization problem. The optimization problem is shown to be NP-hard. We develop a heuristic algorithm based on annotated program graphs, and the algorithm is shown to be practical. Experimental results on an IBM SP-2 show that the selection algorithms are effective in improving the performances of application programs that use sparse data sets. Rong-Guey Chang, Tyng-Ruey Chuang, Jenq Kuen Lee |
ICPP | 3 |
| 1999 | Communication set generations with CSD calculus and expression-rewriting framework
Gwan-Hwan Hwang, Jenq Kuen Lee |
Parallel Comput. | 2 |
| 1998 | Real-Time Gang Schedulings with Workload Models for Parallel ComputersabstractGang scheduling has been shown to be an effective job scheduling policy for parallel computers that combines elements of space sharing and time sharing. We propose new policies to enable gang scheduling to adapt to environments with real-time constraints. Our work, to our best knowledge, is the first work to attempt to address the real-time aspects of gang scheduling. Our system guided by a metric, called "task utilization workload", can schedule both real-time and non-real-time tasks at the same time. We report simulation results with a family of scheduling algorithms based on our proposed metric. Our scheme is designed to be a practical scheme to be used for large scale industrial and commercial parallel systems. Preliminary simulation results also show that our proposed policy is an effective scheme to perform real-time scheduling, while scheduling non-real-time jobs with fairness and good throughput. Jenq Kuen Lee, Chung-Der Lin, Yar-Wen Chang, Wei-Kuan Shih |
ICPADS | 1 |
| 1998 | Efficient Support of Parallel Sparse Computation for Array Intrinsic Functions of Fortran 90abstractFortran 90 provides a rich set of array intrinsic functions. They form a rich source of parallelism and play an increasingly important role in automatic support of data parallel programming. However, there is no such support if these intrinsic functions are applied to sparse data sets. We address this open gap by presenting an efficient library for parallel sparse computations with Fortran 90 array intrinsic operations. Our method provides both compression schemes and distribution schemes on distributed memory environments applicable to higherdimensional sparse arrays. Sparse programs can be expressed concisely using array expressions, and parallelized with the help of our library. Preliminary experimental results on an IBM SP2 workstation cluster show that our approach is promising in supporting efficient sparse matrix computations on both sequential and distributed memory environments. 1 Introduction An increasing number of programming languages, such as APL, Fortran 90, High Perfor... Rong-Guey Chang, Tyng-Ruey Chuang, Jenq Kuen Lee |
International Conference on Supercomputing | 3 |
| 1998 | A Function-Composition Approach to Synthesize Fortran 90 Array Operations
Gwan-Hwan Hwang, Jenq Kuen Lee, Roy Dz-Ching Ju |
J. Parallel Distributed Comput. | 2 |
| 1997 | Data Distribution Analysis and Optimization for Pointer-Based Distributed ProgramsabstractA critical question remains open if the compiler can understand the distribution pattern of pointer-based distributed objects built by application programmers, and perform optimization as effectively as the HPF compiler does with distributed arrays. In this paper, we address this challenging issue. In our work, we first present a parallel progamming model which allows application programmers to build pointer-based distributed objects at application levels. Next we propose a distribution analysis algorithm which can automatically summarize the distribution pattern of pointer-based distributed objects built by application programmers. Our work, to our best knowledge, is the first work to attempt to address this open issue. Our distribution analysis framework employs Feautrier's parametric integer programming as the basic solver, and can always obtain precise distribution information from the class of programs written in our parallel programming model with static control. Experimental results done on a 16-node IBM SP-2 machine show that the compiler with the help of distribution analysis algorithm can significantly improve the performance of pointer-based distributed programs. Jenq Kuen Lee, Dan Ho, Y. C. Chuang |
ICPP | 1 |
| 1997 | Towards Automatic Support of Parallel Sparse Computation in Java with Continuous CompilationabstractWe present a generic matrix class facility in Java and an on-going project for a runtime environment with continuous compilation aiming to support automatic parallelization of sparse computation on distributed environments. Our package comes with a collection of matrix classes with a uniform interface for operations on dense and sparse matrices. These matrix operations are implemented both for sequential and parallel executions on distributed memory environments. In our environment, a program such as the conjugate gradient solver is written by users using high-level generic matrix notations in Java. At runtime the generic notations are mapped to specific implementations. Our approach is particularly useful for optimizing sparse computation for distributed environments because, with the help of profiling information and a cost model, it can automatically select suitable compression and distribution schemes according to access patterns of the programs and non-zero structures of the matrices. Our testbed is currently based on Java and PVM on an IBM SP2 workstation cluster. Preliminary experimental results show that our approach is promising in speeding up sparse matrix computations on distributed memory environments. © 1997 John Wiley & Sons, Ltd. Rong-Guey Chang, Cheng-Wei Chen, Tyng-Ruey Chuang, Jenq Kuen Lee |
Concurr. Pract. Exp. | 4 |
| 1997 | Parallel Array Object I/O Support on Distributed Environments
Jenq Kuen Lee, Ing-Kuen Tsaur, San-Yih Hwang |
J. Parallel Distributed Comput. | 1 |
| 1995 | An Array Operation Synthesis Scheme to Optimize Fortran 90 ProgramsabstractAn increasing number of programming languages, such as Fortran 90 and APL, are providing a rich set of intrinsic array functions and array expressions. These constructs which constitute an important part of data parallel languages provide excellent opportunities for compiler optimizations. In this paper, we present a new approach to combine consecutive data access patterns of array constructs into a composite access function to the source arrays. Our scheme is based on the composition of access functions, which is similar to a composition of mathematic functions. Our new scheme can handle not only data movements of arrays of different numbers of dimensions and segmented array operations but also masked array expressions and multiple sources array operations. As a result, our proposed scheme is the first synthesis scheme which can synthesize Fortran 90 RESHAPE, EOSHIFT, MERGE, and WHERE constructs together. Experimental results show speedups from 1.21 to 2.95 for code fragments from real applications on a Sequent multiprocessor machine by incorporating the proposed optimizations. Gwan-Hwan Hwang, Jenq Kuen Lee, Roy Dz-Ching Ju |
PPoPP | 2 |
| 1993 | The Xthreads library: Design, implementation, and applicationsabstractThe purpose of the Xthreads library is to provide a cheap concurrent programming environment. The design of the Xthreads library is patterned after Xinu, a small and elegant operating system in which all processes share a single address space and hence enjoy reduced overheads in process creation, interprocess communication, and so on. Our approach is to map the Xinu process structure into the Xthreads thread structure in a Unix-like process. Easy extensions and modifications to the Xthreads library are a major objective, accomplished through modularity and layering. We ported Xthreads to the nCUBE2, iPSC860 and RS6000 computers. This paper describes the library, our experiences with its design and implementation, the early performance measurements, and its applicability to simulation modeling.> Janche Sang, Felipe Knop, Vernon Rego, Jenq Kuen Lee, Chung-Ta King |
COMPSAC | 4 |
| 1991 | Object oriented parallel programming: experiments and resultsabstractWe present cm object-oriented, parallel programming paradigm, called the distributed collection model and an experimental language PC++ based on the model.In the distributed collection model, programmers can describe the data distribution of elements among processors to utilize memoy locality and a collection construct is employed to build distributed structures.The model also supports the express of massive parallelism and a new mechanism for building hierarchies of abstractions.Our experiences with application programs in the PC++ programming environments as well as performance results are also described in the paper. Jenq Kuen Lee, Dennis Gannon |
SC | 1 |