I-Hsin Chung

dblp:74/1723 · DBLP profile ↗
← Back
51ranked-venue papers
12as first author
12since 2021 · last 2026
0000-0003-4555-9257ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 34 · 11 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Data-Driven Lipschitz Continuity: A Cost-Effective Approach to Improve Adversarial Robustness
abstract
As deep neural networks (DNNs) are increasingly deployed in sensitive applications, ensuring their security and robustness has become critical. A major threat to DNNs arises from adversarial attacks, where small input perturbations can lead to incorrect predictions. Recent advances in adversarial training improve robustness by incorporating additional examples from external datasets or generative models. However, these methods often incur high computational costs, limiting their practicality and hindering real-world deployment. In this paper, we propose a cost-efficient alternative based on Lipschitz continuity that achieves robustness comparable to models trained with extensive supplementary data. Unlike conventional adversarial training, our method requires only a single pass over the dataset without gradient estimation, making it highly efficient. Furthermore, our method can integrate seamlessly with existing adversarial training frameworks and enhances the robustness of models without requiring extra generative data. Experimental results show that our approach not only reduces computational overhead but also maintains or improves the defensive capabilities of robust neural networks. This work opens a promising direction for developing practical, scalable defenses against adversarial attacks.
Erh-Chung Chen, I-Hsin Chung, Che-Rung Lee
WACV3
2025 Vela: A Virtualized LLM Training System with GPU Direct RoCE
abstract
Vela is a cloud-native system designed for LLM training workloads built using off-the-shelf hardware, Linux KVM-based virtualization, and a virtualized RDMA over Converged Ethernet (RoCE) network. Vela virtual machines (VMs) support peer-to-peer DMA between the GPUs and SRIOV-based network interface. In this paper, we share Vela's key architectural aspects with details from an NVIDIA A100 GPU-based deployment in one of the IBM Cloud data centers. Throughout the paper, we share insights and experiences from designing, building, and operating the system over a ~2.5 year timeframe to highlight the capabilities of readily available software and hardware technologies and the improvement opportunities for future AI systems, thereby making AI infrastructure more accessible to a broader community. As we evaluated the system for performance at ~1500 GPU scale, we achieved ~80% of the ideal throughput while training a 50 billion parameter decoder model using model parallelism, and ~70% per GPU FLOPS compared to a single VM with the High-Performance Linpack benchmark.
Apoorve Mohan, Robert Walkup, Bengi Karaçali, Ming-Hung Chen, Abdullah Kayi, Liran Schour, Shweta Salaria, Sophia Wen, I-Hsin Chung, Abdul Alim, Constantinos Evangelinos, Lixiang Luo, Marc Dombrowa, Laurent Schares, Ali Sydney, Pavlos Maniotis, Sandhya Koteshwara, Brent Tang, Joel Belog, Rei Odaira, Vasily Tarasov, Eran Gampel, Drew Thorstensen, Talia Gershon, Seetharami Seelam
ASPLOS (2)9
2025 Fast Malicious Packets Inspection Framework Using Converged Accelerator
Chuan-Ming Ou, Yong-Xuan Huang, Ming-Hung Chen, I-Hsin Chung, Jerry Chou 0001
HPC Asia4
2025 PCIe Bandwidth-Aware Scheduling for Multi-Instance GPUs
Yan-Mei Tang, Wei-Fang Sun, Hsu-Tzu Ting, Ming-Hung Chen, I-Hsin Chung, Jerry Chou 0001
HPC Asia5
2025 Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
abstract
We present the design and implementation of a new lifetime-aware tensor offloading framework for GPU memory expansion using low-cost PCIe-based solid-state drives (SSDs). Our framework, TERAIO, is developed explicitly for large language model (LLM) training with multiple GPUs and multiple SSDs. Its design is driven by our observation that the active tensors take only a small fraction (1.7% on average) of allocated GPU memory in each LLM training iteration, the inactive tensors are usually large and will not be used for a long period of time, creating ample opportunities for offloading/prefetching tensors to/from slow SSDs without stalling the GPU training process. TERAIO accurately estimates the lifetime (active period of time in GPU memory) of each tensor with the profiling of the first few iterations in the training process. With the tensor lifetime analysis, TERAIO will generate an optimized tensor offloading/prefetching plan and integrate it into the compiled LLM program via PyTorch. TERAIO has a runtime tensor migration engine to execute the offloading/prefetching plan via GPUDirect storage, which allows direct tensor migration between GPUs and SSDs for alleviating the CPU bottleneck and maximizing the SSD bandwidth utilization. In comparison with state-of-the-art studies such as ZeRO-Offload and ZeRO-Infinity, we show that TERAIO improves the training performance of various LLMs by 1.47× on average, and achieves 80.7% of the ideal performance assuming unlimited GPU memory.
Yirui Eric Zhou, Apoorve Mohan, I-Hsin Chung, Seetharami Seelam, Jian Huang 0006
NeurIPS5
2024 Latency Attack Resilience in Object Detectors: Insights from Computing Architecture
Erh-Chung Chen, I-Hsin Chung, Che-Rung Lee
ACCV (8)3
2024 Overload: Latency Attacks on Object Detection for Edge Devices
abstract
Nowadays, the deployment of deep learning-based applications is an essential task owing to the increasing demands on intelligent services. In this paper, we investigate latency attacks on deep learning applications. Unlike common adversarial attacks for misclassification, the goal of latency attacks is to increase the inference time, which may stop applications from responding to the requests within a reasonable time. This kind of attack is ubiquitous for various applications, and we use object detection to demonstrate how such kind of attacks work. We also design a framework named Overload to generate latency attacks at scale. Our method is based on a newly formulated optimization problem and a novel technique, called spatial attention. This attack serves to escalate the required computing costs during the inference time, consequently leading to an extended inference time for object detection. It presents a significant threat, especially to systems with limited computing resources. We conducted experiments using YOLOv5 models on Nvidia NX. Compared to existing methods, our method is simpler and more effective. The experimental results show that with latency attacks, the inference time of a single image can be increased ten times longer in reference to the normal setting. Moreover, our findings pose a potential new threat to all object detection tasks requiring non-maximum suppression (NMS), as our attack is NMS-agnostic.
Erh-Chung Chen, I-Hsin Chung, Che-Rung Lee
CVPR3
2023 Enabling Scalability in the Cloud for Scientific Workflows: An Earth Science Use Case
abstract
Scientific discovery increasingly relies on interoperable, multimodular workflows generating intermediate data. The complexity of managing intermediate data may cause performance losses or unexpected costs. This paper defines an approach to composing these scientific workflows on cloud services, focusing on workflow data orchestration, management, and scalability. We demonstrate the effectiveness of our approach with the SOMOSPIE scientific workflow that deploys machine learning (ML) models to predict high-resolution soil moisture using an HPC service (LSF) and an open-source cloud-native service (K8s) and object storage. Our approach enables scientists to scale from coarse-grained to fine-grained resolution and from a small to a larger region of interest. Using our empirical observations, we generate a cost model for the execution of workflows with hidden intermediate data on cloud services.
Paula Olaya, Jakob Lüttgau, Camila Roa, Ricardo M. Llamas, Rodrigo Vargas, Sophia Wen, I-Hsin Chung, Seetharami R. Seelam, Yoonho Park, Jay F. Lofstead, Michela Taufer
CLOUD7
2023 GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System Architecture
abstract
Graphics Processing Units (GPUs) have traditionally relied on the host CPU to initiate access to the data storage. This approach is well-suited for GPU applications with known data access patterns that enable partitioning of their dataset to be processed in a pipelined fashion in the GPU. However, emerging applications such as graph and data analytics, recommender systems, or graph neural networks, require fine-grained, data-dependent access to storage. CPU initiation of storage access is unsuitable for these applications due to high CPU-GPU synchronization overheads, I/O traffic amplification, and long CPU processing latencies. GPU-initiated storage removes these overheads from the storage control path and, thus, can potentially support these applications at much higher speed. However, there is a lack of systems architecture and software stack that enable efficient GPU-initiated storage access. This work presents a novel system architecture, BaM, that fills this gap. BaM features a fine-grained software cache to coalesce data storage requests while minimizing I/O traffic amplification. This software cache communicates with the storage system via high-throughput queues that enable the massive number of concurrent threads in modern GPUs to make I/O requests at a high rate to fully utilize the storage devices and the system interconnect. Experimental results show that BaM delivers 1.0x and 1.49x end-to-end speed up for BFS and CC graph analytics benchmarks while reducing hardware costs by up to 21.7x over accessing the graph data from the host memory. Furthermore, BaM speeds up data-analytics workloads by 5.3x over CPU-initiated storage access on the same hardware.
Zaid Qureshi, Vikram S. Mailthody, Isaac Gelado, Seungwon Min, Amna Masood, Jeongmin Brian Park, Jinjun Xiong, Chris J. Newburn, Dmitri Vainbrand, I-Hsin Chung, Michael Garland, William J. Dally, Wen-Mei W. Hwu
ASPLOS (2)10
2022 NVMe Virtualization for Cloud Virtual Machines
abstract
Public clouds are rapidly moving to support Non-Volatile Memory Express (NVMe) based storage to meet the ever-increasing I/O throughput and latency demands of modern workloads. They provide NVMe storage through virtual machines (VMs) where multiple VMs running on a host may share a physical NVMe device. The virtualization method used to share the NVMe capability has important performance, usability and security implications. In this paper, we propose three NVMe storage virtualization methods: PCI device passthrough, virtual block device method, and Storage Performance Development Kit (SPDK) virtual host target method. We evaluate these virtualization methods in terms of performance, scalability, CPU overhead, technology maturity, security, and availability to use one or more of these methods in IBM public cloud.
Lixiang Luo, I-Hsin Chung, Seetharami R. Seelam, Ming-Hung Chen, Yun Joon Soh
ICPE2
2021 A Deep Reinforcement Learning Method for Solving Task Mapping Problems with Dynamic Traffic on Parallel Systems
abstract
Efficient mapping of application communication patterns to the network topology is a critical problem for optimizing the performance of communication bound applications on parallel computing systems. The problem has been extensively studied in the past, but they mostly formulate the problem as finding an isomorphic mapping between two static graphs with edges annotated by traffic volume and network bandwidth. But in practice, the network performance is difficult to be accurately estimated, and communication patterns are often changing over time and not easily obtained. Therefore, this work proposes a deep reinforcement learning (DRL) approach to explore better task mappings by utilizing the performance prediction and runtime communication behaviors provided from a simulator to learn an efficient task mapping algorithm. We extensively evaluated our approach using both synthetic and real applications with varied communication patterns on Torus and Dragonfly networks. Compared with several existing approaches from literature and software library, our proposed approach found task mappings that consistently achieved comparable or better application performance. Especially for a real application, the average improvement of our approach on Torus and Dragonfly networks are 11% and 16%, respectively. In comparison, the average improvements of other approaches are all less than 6%.
YuCheng Wang, Jerry Chou 0001, I-Hsin Chung
HPC Asia3
2021 TEMPI: An Interposed MPI Library with a Canonical Representation of CUDA-aware Datatypes
abstract
MPI derived datatypes are an abstraction that simplifies handling of non-contiguous data in MPI applications. These datatypes are recursively constructed at runtime from primitive Named Types defined in the MPI standard. More recently, the development and deployment of CUDA-aware MPI implementations has encouraged the transition of distributed high-performance MPI codes to use GPUs. Such implementations allow MPI functions to directly operate on GPU buffers, easing integration of GPU compute into MPI codes. This work first presents a novel datatype handling strategy for nested strided datatypes, which finds a middle ground between the specialized or generic handling in prior work. This work also shows that the performance characteristics of non-contiguous data handling can be modeled with empirical system measurements, and used to transparently improve MPI_Send/Recv latency. Finally, despite substantial attention to non-contiguous GPU data and CUDA-aware MPI implementations, good performance cannot be taken for granted. This work demonstrates its contributions through an MPI interposer library, TEMPI. TEMPI can be used with existing MPI deployments without system or application changes. Ultimately, the interposed-library model of this work demonstrates MPI_Pack speedup of up to 242000x and MPI_Send speedup of up to 59000x compared to the MPI implementation deployed on a leadership-class supercomputer. This yields speedup of more than 917x in a 3D halo exchange with 3072 processes.
Carl Pearson, Kun Wu 0002, I-Hsin Chung, Jinjun Xiong, Wen-Mei W. Hwu
HPDC3
2020 ECS2: A Fast Erasure Coding Library for GPU-Accelerated Storage Systems with Parallel & Direct IO
abstract
As data volume keeps increasing at a rapid rate, there is an urgent need for large, reliable, and cost-effective storage systems. Erasure coding has drawn increasing attention because of its ability to ensure data reliability with higher storage efficiency, and it has been widely adopted in many distributed and large-scale storage systems, such as Azure cloud storage and HDFS. However, the storage efficiency of erasure code comes at the price of higher computing complexity. While many studies have shown the coding computations can be significantly accelerated using GPU, the overhead of data transfer between storage devices and GPUs become a new performance bottleneck. In this work, we designed and implemented, ECS2, a fast erasure coding library on GPU -accelerated storage to let users enhance their data protection with transparent IO performance and file system like programming interface. By taking advantage of the latest GPUDirect technology supported on Nvidia GPU, our library is able to bypass CPU and host memory copy from the IO path, so that both the computing and IO overhead from coding can be minimized. Using synthetic IO workload based on real storage system trace, we show that the IO latency can be reduced by 10% ~ 20% with GPUDirect technology, and the overall IO throughput of a storage system can be improved up to 70%.
ChanJung Chang, Jerry Chou 0001, Yu-Ching Chou, I-Hsin Chung
CLUSTER4
2019 Towards VR/AR Multimedia Content Multicast over Wireless LAN
abstract
Virtual reality (VR) and augmented reality (AR) are expected to change people's daily life in the future, but it still faces many challenges. New advanced coding techniques like depth image based rendering (DIBR) mitigate the need to complete content by synthesizing the scene with nearby views. However, due to the limitation of wireless spectrum, delivering VR/AR content to many people doing some activities (e.g., watching sport in a stadium or seeing an exhibition in a museum) in mobile hotspot areas is still very challenging in terms of available bandwidth. In this work, we analyze how to multicast multi-view VR/AR 3D multimedia streams over multiple wireless access points (APs) to multiple users, and then propose a multi-view allocation (MVA) algorithm to improve the delivered content quality. The simulation results show that our MVA algorithm can achieve about 90% performance of the optimal solution and outperform non-DIBR and SSA model, by 48% and 25%, respectively.
Ming-Hung Chen, Kai-Wen Hu, I-Hsin Chung, Cheng-Fu Chou
CCNC3
2019 Temporal-based Load Adaptive SDN Controller Failover Mechanism
abstract
Recently distributed multiple SDN controller architecture is proposed to improve the scalability of SDN architecture. Unfortunately, the state-of-art SDN controller failover mechanism, i.e., switch leader competition, did not consider the characteristics of data center network and the time cost, resource usage, and load balancing after failure. In this paper, we propose a temporal-based load adaptive SDN controller failover mechanism (TLACF) based on the temporal pattern of traffic. The results demonstrate the proposed TLACF can save up to 46% time cost of failover and 66% memory usage for state backup.
Ming-Hung Chen, Zhi-Qiang Zhong, I-Hsin Chung, Cheng-Fu Chou
CCNC3
2019 Evaluating Characteristics of CUDA Communication Primitives on High-Bandwidth Interconnects
abstract
Data-intensive applications such as machine learning and analytics have created a demand for faster interconnects to avert the memory bandwidth wall and allow GPUs to be effectively leveraged for lower compute intensity tasks. This has resulted in wide adoption of heterogeneous systems with varying underlying interconnects, and has delegated the task of understanding and copying data to the system or application developer. No longer is a malloc followed by memcpy the only or dominating modality of data transfer; application developers are faced with additional options such as unified memory and zero-copy memory. Data transfer performance on these systems is now impacted by many factors including data transfer modality, system interconnect hardware details, CPU caching state, CPU power management state, driver policies, virtual memory paging efficiency, and data placement.
Carl Pearson, Abdul Dakkak, Sarah Hashash, Cheng Li 0014, I-Hsin Chung, Jinjun Xiong, Wen-Mei W. Hwu
ICPE5
2018 FlexProtect: A SDN-based DDoS Attack Protection Architecture for Multi-tenant Data Centers
abstract
With the recent advances in software-defined networking (SDN), the multi-tenant data centers provide more efficient and flexible cloud platform to their subscribers. However, as the number, scale, and diversity of distributed denial-of-service (DDoS) attack is dramatically escalated in recent years, the availability of those platforms is still under risk. We note that the state-of-art DDoS protection architectures did not fully utilize the potential of SDN and network function virtualization (NFV) to mitigate the impact of attack traffic on data center network. Therefore, in this paper, we exploit the flexibility of SDN and NFV to propose FlexProtect, a flexible distributed DDoS protection architecture for multi-tenant data centers. In FlexProtect, the detection virtual network functions (VNFs) are placed near the service provider and the defense VNFs are placed near the edge routers for effectively detection and avoid internal bandwidth consumption, respectively. Based on the architecture, we then propose FP-SYN, an anti-spoofing SYN flood protection mechanism. The emulation and simulation results with real-world data demonstrates that, compared with the traditional approach, the proposed architecture can significantly reduce 46% of the additional routing path and save 60% internal bandwidth consumption. Moreover, the proposed detection mechanism for anti-spoofing can achieve 98% accuracy.
Ming-Hung Chen, Jyun-Yan Ciou, I-Hsin Chung, Cheng-Fu Chou
HPC Asia3
2018 Towards a Composable Computer System
abstract
The recent advancement of technology in both software and hardware enables us to revisit the concept of the composable architecture in the system design. The composable system design provides flexibility to serve a variety of workloads. The system offers a dynamic co-design platform that allows experiments and measurements in a controlled environment. This speeds up the system design and software evolution. It also decouples the lifecycles of components. The design consideration includes adopting available technology with the understanding of application characteristics. With the flexibility, we show the design has the potential to be the infrastructure of both cloud computing and HPC architecture serving a variety of workloads.
I-Hsin Chung, Bülent Abali, Paul Crumley
HPC Asia1
2018 Towards a Single-Host Many-GPU System
abstract
As computation-intensive tasks such as deep learning and big data analysis take advantage of GPU based accelerators, the interconnection links may become a bottleneck. In this paper, we investigate the upcoming performance bottleneck of multi-accelerator systems, as the number of accelerators equipped with single host grows. We instrumented the host PCIe fabric to measure the data transfer and compared it with the measurements from the software tool. It shows how the data transfer (P2P) helps to avoid the bottleneck on the interconnection links, but multi-GPU performance does not scale up as expected due to the control messages. We quantify the impact of host control messages with suggestions to remedy scalability bottlenecks. We also implement the proposed strategy on Lulesh to validate the concept. The result shows our strategy can save 59.86% time cost of the kernel and 13.32% PCIe H2D payload.
Ming-Hung Chen, I-Hsin Chung, Bülent Abali, Paul Crumley
SBAC-PAD2
2017 Parallel Deep Neural Network Training for Big Data on Blue Gene/Q
abstract
Deep Neural Networks (DNNs) have recently been shown to significantly outperform existing machine learning techniques in several pattern recognition tasks. DNNs are the state-of-the-art models used in image recognition, object detection, classification and tracking, and speech and language processing applications. The biggest drawback to DNNs has been the enormous cost in computation and time taken to train the parameters of the networks-often a tenfold increase relative to conventional technologies. Such training time costs can be mitigated by the application of parallel computing algorithms and architectures. However, these algorithms often run into difficulties because of the cost of inter-processor communication bottlenecks. In this paper, we describe how to enable Parallel Deep Neural Network Training on the IBM Blue Gene/Q (BG/Q) computer system. Specifically, we explore DNN training using the data-parallel Hessian-free 2nd order optimization algorithm. Such an algorithm is particularly well-suited to parallelization across a large set of loosely coupled processors. BG/Q, with its excellent inter-processor communication characteristics, is an ideal match for this type of algorithm. The paper discusses how issues regarding programming model and data-dependent imbalances are addressed. Results on large-scale speech tasks show that the performance on BG/Q scales linearly up to 4,096 processes with no loss in accuracy. This allows us to train neural networks using billions of training examples in a few hours.
I-Hsin Chung, Tara N. Sainath, Bhuvana Ramabhadran, Michael Picheny, John A. Gunnels, Vernon Austel, Upendra V. Chaudhari, Brian Kingsbury
IEEE Trans. Parallel Distributed Syst.1
2014 Parallel deep neural network training for LVCSR tasks using blue gene/Q
abstract
While Deep Neural Networks (DNNs) have achieved tremendous success for LVCSR tasks, training these networks is slow. To date, the most common approach to train DNNs is via stochastic gradient descent (SGD), serially on a single GPU machine. Serial training, coupled with the large number of training parameters and speech data set sizes, makes DNN training very slow for LVCSR tasks. While 2nd order, data-parallel methods have also been explored, these methods are not always faster on CPU clusters due to the large communication cost between processors. In this work, we explore using a specialized hardware/software approach, utilizing a Blue Gene/Q (BG/Q) system, which has thousands of processors and excellent interprocessor communication. We explore using the 2nd order Hessian-free (HF) algorithm for DNN training with BG/Q, for both cross-entropy and sequence training of DNNs. Results on three LVCSR tasks indicate that using HF with BG/Q offers up to an 11x speedup, as well as an improved word error rate (WER), compared to SGD on a GPU.
Tara N. Sainath, I-Hsin Chung, Bhuvana Ramabhadran, Michael Picheny, John A. Gunnels, Brian Kingsbury, George Saon, Vernon Austel, Upendra V. Chaudhari
INTERSPEECH2
2014 Parallel Deep Neural Network Training for Big Data on Blue Gene/Q
abstract
Deep Neural Networks (DNNs) have recently been shown to significantly outperform existing machine learning techniques in several pattern recognition tasks. DNNs are the state-of-the-art models used in image recognition, object detection, classification and tracking, and speech and language processing applications. The biggest drawback to DNNs has been the enormous cost in computation and time taken to train the parameters of the networks - often a tenfold increase relative to conventional technologies. Such training time costs can be mitigated by the application of parallel computing algorithms and architectures. However, these algorithms often run into difficulties because of the cost of inter-processor communication bottlenecks. In this paper, we describe how to enable Parallel Deep Neural Network Training on the IBM Blue Gene/Q (BG/Q) computer system. Specifically, we explore DNN training using the data parallel Hessian-free 2nd order optimization algorithm. Such an algorithm is particularly well-suited to parallelization across a large set of loosely coupled processors. BG/Q, with its excellent inter-processor communication characteristics, is an ideal match for this type of algorithm. The paper discusses how issues regarding programming model and data-dependent imbalances are addressed. Results on large-scale speech tasks show that the performance on BG/Q scales linearly up to 4096 processes with no loss in accuracy. This allows us to train neural networks using billions of training examples in a few hours.
I-Hsin Chung, Tara N. Sainath, Bhuvana Ramabhadran, Michael Picheny, John A. Gunnels, Vernon Austel, Upendra V. Chaudhari, Brian Kingsbury
SC1
2014 Improving GPU Memory Performancewith Artificial Barrier Synchronization
abstract
Barrier synchronization, an essential mechanism for a block of threads to guard data consistency, is regarded as a threat to performance. This study, however, provides a different viewpoint for barrier synchronization on GPUs: adding barrier synchronization, even when functionally unnecessary, can improve the performance of some memory-intensive applications. We explain this phenomenon using a memory contention model in which artificial barrier synchronization helps reduce memory contention and preserve data access locality. To yield practical applications, we identify a program pattern: artificial barrier synchronization can be used to synchronize the memory accesses when the data locality among threads is violated. Empirical results from three real-world applications demonstrate that artificial barrier synchronization can increase performance by 10 to 20 percent.
Shih-Hsiang Lo, Che-Rung Lee, Quey-Liang Kao, I-Hsin Chung, Yeh-Ching Chung
IEEE Trans. Parallel Distributed Syst.4
2013 TLA: Temporal look-ahead processor allocation method for heterogeneous multi-cluster systems
Po-Chi Shih, Kuo-Chan Huang, Che-Rung Lee, I-Hsin Chung, Yeh-Ching Chung
J. Parallel Distributed Comput.4
2012 GPU Performance Enhancement via Communication Cost Reduction: Case Studies of Radix Sort and WSN Relay Node Placement Problem
abstract
As the computational power of Graphics Processing Unit (GPU) increases, data transmission becomes the major performance bottleneck. In this study, we investigate two techniques, data streaming and data compression, to reduce the communication cost on GPU. Data streaming enables overlap of communication and computation, whereas data compression reduces the data size transferred among different memory spaces. Although both techniques increase computation cost, overall performance can still be enhanced by reducing communication cost. We demonstrate the effectiveness of the two techniques via two case studies: radix sort and 3-star, a deployment algorithm in wireless sensor networks. For radix sort, a new algorithm, which mixes MSD and LSD algorithms and employs data streaming, is presented. Its performance is 25% faster than the fastest GPU radix sort implementation currently available in the public domain. For the 3-star algorithm, the speed increases several hundreds of times faster than that obtained by the CPU code. The data streaming and data compression, which is a hybrid CPU-GPU algorithm, provide an additional 54% performance improvement to the GPU implementation. Data compression not only reduces communication cost, but also improves the computation time, by which further performance enhancement can be achieved.
Che-Rung Lee, Shih-Hsiang Lo, Nan-Hsi Chen, Yeh-Ching Chung, I-Hsin Chung
CCGRID5
2012 An Efficient Framework for Multi-dimensional Tuning of High Performance Computing Applications
abstract
Deploying an application onto a target platform for high performance oftentimes demands manual tuning by experts. As machine architecture gets increasingly complex, tuning becomes even more challenging and calls for systematic approaches. In our earlier work we presented a prototype that combines efficiently expert knowledge, static analysis, and runtime observation for bottleneck detection, and employs refactoring and compiler feedback for mitigation. In this study, we develop a software tool that facilitates \emph{fast} searching of bottlenecks and effective mitigation of problems from major dimensions of computing (e.g., computation, communication, and I/O). The impact of our approach is demonstrated by the tuning of the LBMHD code and a Poisson solver code, representing traditional scientific codes, and a graph analysis code in UPC, representing emerging programming paradigms. In the experiments, our framework detects with a single run of the application intricate bottlenecks of memory access, I/O, and communication. Moreover, the automated solution implementation yields significant overall performance improvement on the target platforms. The improvement for LBMHD is up to 45\%, and the speedup for the UPC code is up to 5. These results suggest that our approach is a concrete step towards systematic tuning of high performance computing applications.
Guojing Cong, Hui-Fang Wen, I-Hsin Chung, David J. Klepacki, Hiroki Murata, Yasushi Negishi
IPDPS3
2012 Application data prefetching on the IBM blue gene/Q supercomputer
abstract
Memory access latency is often a crucial performance limitation for high performance computing. Prefetching is one of the strategies used by system designers to bridge the processor-memory gap. This paper describes a new innovative list prefetching feature introduced in the IBM Blue Gene/Q supercomputer. The list prefetcher records the L1 cache miss addresses and prefetches them in the next iteration. The evaluation shows this list prefetching mechanism reduces data fetching time when L1 cache misses happen and improves the performance for high performance computing applications with repeating nonuniform memory access patterns. Its performance is compatible with classic stream prefetcher when properly configured.
I-Hsin Chung, Changhoan Kim, Hui-Fang Wen, Guojing Cong
SC1
2012 A Systematic Approach toward Automated Performance Analysis and Tuning
abstract
High productivity is critical in harnessing the power of high-performance computing systems to solve science and engineering problems. It is a challenge to bridge the gap between the hardware complexity and the software limitations. Despite significant progress in programming language, compiler, and performance tools, tuning an application remains largely a manual task, and is done mostly by experts. In this paper, we propose a systematic approach toward automated performance analysis and tuning that we expect to improve the productivity of performance debugging significantly. Our approach seeks to build a framework that facilitates the combination of expert knowledge, compiler techniques, and performance research for performance diagnosis and solution discovery. With our framework, once a diagnosis and tuning strategy has been developed, it can be stored in an open and extensible database and thus be reused in the future. We demonstrate the effectiveness of our approach through the automated performance analysis and tuning of two scientific applications. We show that the tuning process is highly automated, and the performance improvement is significant.
Guojing Cong, I-Hsin Chung, Hui-Fang Wen, David J. Klepacki, Hiroki Murata, Yasushi Negishi, Takao Moriyama
IEEE Trans. Parallel Distributed Syst.2
2011 A Parallel Rectangle Intersection Algorithm on GPU+CPU
abstract
In this paper, we investigate efficient algorithms and implementations using GPU plus CPU to solve the rectangle intersection problem on a plane. The problem is to report all intersecting pairs of iso-oriented rectangles, whose parallelization on GPUs poses two major computational challenges: data partition and the massive output. The algorithm we presented is called PRI-GC, Parallel Rectangle Intersection algorithm on GPU+CPU, which consists of two phases: mapping and intersection-checking. In the mapping phase, rectangles are hashed into different subspaces (called cells) to reduce the unnecessary intersection checking for far-apart rectangles. In the intersection-checking phase, pairs of rectangles within the same cell are examined in parallel, and the intersecting pairs of rectangles are reported. Several optimization techniques, including rectangles re-ordering, output data compressing/encoding, and the execution overlapping of GPU and CPU, are applied to enhance the performance. We had evaluated the performance of PRI-GC and the result shows over 30x speedup against two well-implemented sequential algorithms on single CPU. The effectiveness of each optimization technique for this problem was evaluated as well. Several parameters, including different degrees of rectangle coverage, different block sizes, and different cell sizes, were also experimented to explore their influences on the performance of PRI-GC.
Shih-Hsiang Lo, Che-Rung Lee, Yeh-Ching Chung, I-Hsin Chung
CCGRID4
2011 A Performance Goal Oriented Processor Allocation Technique for Centralized Heterogeneous Multi-cluster Environments
abstract
This paper proposes a processor allocation technique named temporal look-ahead processor allocation (TLPA) that makes allocation decision by evaluating the allocation effects on subsequent jobs in the waiting queue. TLPA has two strengths. First, it takes multiple performance factors into account when making allocation decision. Second, it can be used to optimize different performance metrics. To evaluate the performance of TLPA, we compare TLPA with best-fit and fastest-first algorithms. Simulation results show that TLPA has up to 32.75% performance improvement over conventional processor allocation algorithms in terms of average turnaround time in various system configurations.
Po-Chi Shih, Kuo-Chan Huang, Che-Rung Lee, I-Hsin Chung, Yeh-Ching Chung
CCGRID4
2011 Scalable Communication-Aware Task Mapping Algorithms for Interconnected Multicore Systems
abstract
Communication-aware task mapping algorithms, which map parallel tasks onto processing nodes according to the communication patterns of applications, are essential to reduce the communication time in modern high performance computing. In this paper, we design algorithms specifically for interconnected multicore systems, whose architectural property, namely small number of cores per node, large number of nodes, and large performance gap between the communication within a multicore and among multicores, had brought new challenges and opportunities to the mapping problem. Let k be the number of cores per multicore and n be the number of tasks. We consider the practical case that k ≪ n for k = 2,4, and 6. The designed algorithms are optimal for the mapping measurement, called Maximum Interconnective Message Size (MIMS), and of time complexity merely O(m log m) for m communication pairs. Thus, they are highly scalable for large applications. We had experimented the algorithms on the IBM Blue Gene/P system for two synthetic benchmarks and two applications. The results show good communication performance improvement.
I-Hsin Chung, Che-Rung Lee, Jiazheng Zhou, Yeh-Ching Chung
HPCC1
2010 Automated mapping of regular communication graphs on mesh interconnects
abstract
Network contention has a significantly adverse effect on the performance of parallel applications with increasing size of parallel machines. Machines of the petascale era are forcing application developers to map tasks intelligently to job partitions to achieve the best performance possible. This paper presents a framework for automated mapping of parallel applications with regular communication graphs to two and three dimensional mesh and torus networks. This framework will save much effort on the part of application developers to generate mappings for their individual applications. One component of the framework is a process topology analyzer to find regular patterns and if found, to determine the dimensions of the communication graphs of applications. The other component is a suite of heuristic techniques for mapping 2D object grids to 2D and 3D processor meshes. The framework chooses the best heuristic from the suite for a given object grid and processor mesh pair based on the hop-bytes metric. We show performance improvements using the framework, for a 2D Stencil benchmark in MPI and the Weather Research and Forecasting model running on the IBM Blue Gene/P. We also compare our algorithms with others discussed in literature.
Abhinav Bhatele, Gagan Raj Gupta 0001, Laxmikant V. Kalé, I-Hsin Chung
HiPC4
2010 Parallelization of DQMC simulation for strongly correlated electron systems
abstract
Determinant Quantum Monte Carlo (DQMC) simulation has been widely used to reveal macroscopic properties of strong correlated materials. However, parallelization of the DQMC simulation is extremely challenging duo to the serial nature of underlying Markov chain and numerical stability issues. We extend previous work with novelty by presenting a hybrid granularity parallelization (HGP) scheme that combines algorithmic and implementation techniques to speed up the DQMC simulation. From coarse-grained parallel Markov chain and task decompositions to fine-grained parallelization methods for matrix computations and Green's function calculations, the HGP scheme explores the parallelism on different levels and maps the underlying algorithms onto different computational components that are suitable for modern high performance heterogeneous computer systems. Practical techniques, such as communication and computation overlapping, message compression and load balancing are also considered in the proposed HGP scheme. We have implemented the DQMC simulation with the HGP scheme on an IBM Blue Gene/P system. The effectiveness of the new scheme is demonstrated through both theoretical analysis and performance results. Experiments have shown over a factor of 80 speedups on an IBM Blue Gene/P system with 1,014 computational processors.
Che-Rung Lee, I-Hsin Chung, Zhaojun Bai
IPDPS2
2010 Masking I/O latency using application level I/O caching and prefetching on Blue Gene systems
abstract
In this paper, we present an application-level I/O caching, prefetching, asynchronous system to hide access latency experienced by HPC applications. Our solution of user controllable caching and prefetching system maintains a file-IO cache in the user space of the application, analyzes the I/O access patterns, prefetches requests, and performs write-back of dirty data to storage asynchronously. So each time the application needs the data it does not have to pay the full I/O latency penalty in going to the storage and getting the required data. We have implemented this caching and asynchronous access system on the Blue Gene (BG/L and BG/P) systems. We present experimental results with NAS BT, MADbench, and WRF benchmarks. The results on BG/P system demonstrate that our method hides access latency, enhances application I/O access time by as much as 100%, and improves WRF execution time over 10%.
Seetharami R. Seelam, I-Hsin Chung, John Bauer, Hui-Fang Wen
IPDPS2
2010 Workload performance characterization of DARPA HPCS benchmarks
abstract
Abstract It is critical to understand the workload characteristics and resource usage patterns of available applications to guide the design and development of hardware and software stacks of future machines. In this article, we analyze the workload performance characteristics of three large‐scale DARPA HPCS benchmarks: Hybrid Coordinate Ocean Model, Parallel Ocean Program, and Lattice Boltzemann Magneto‐Hydrodynamics Code while executing on IBM Power5+ processor machines. Our analysis is focused on the CPU/memory performance using Cycles Per Instruction (CPI) model and multiprocess communication performance using MPI traces. For each benchmark, we provide a high‐level performance analysis followed by the hotspot analysis for selected input parameters. Then we present a detailed workload performance characterization using CPI model with data from a unique set of performance counters available on the Power5+ processor system. From communication performance analysis, we describe the sources of load imbalances in the applications and identify the potential impediments to the scalability of the applications under large processor counts. We identify several sources of performance problems that are potential bottlenecks and discuss methods to ameliorate them. We also present a comparative analysis of these benchmarks to summarize the similarities and differences in their performance characteristics. Copyright © 2009 John Wiley & Sons, Ltd.
Seetharami R. Seelam, I-Hsin Chung, Guojing Cong, Hui-Fang Wen, David J. Klepacki
Concurr. Comput. Pract. Exp.2
2009 A Holistic Approach towards Automated Performance Analysis and Tuning
Guojing Cong, I-Hsin Chung, Hui-Fang Wen, David J. Klepacki, Hiroki Murata, Yasushi Negishi, Takao Moriyama
Euro-Par2
2009 Tools for scalable performance analysis on Petascale systems
abstract
Tools are becoming increasingly important to efficiently utilize the computing power available in contemporary large scale systems. The drastic increase in the size and the complexity of systems require tools to be scalable while producing meaning full and easily digestible information that may help the user pin-point problems at scale. The goal of this tutorial is to introduce some state-of-the-art performance tools from three different organizations to a diverse audience group. Together these tools provide a broad spectrum of capabilities necessary to analyze the performance of scientific and engineering applications on a variety of large and small scale systems. These tools include: • IBM High Performance Computing Toolkit: The IBM High Performance Computing Toolkit is a suite of performance-related tools and libraries to assist in application tuning. This toolkit is an integrated environment for performance analysis of sequential and parallel applications using the MPI and OpenMP paradigms. Scientists can collect rich performance data from selected parts of an execution, digest the data at a very high level, and plan for improvements within a single unified interface. It provides a common framework for IBM's mid-range server offerings, including pSeries and eSeries servers and Blue Gene systems, on both AIX and Linux. More information cab be found here: http://domino.research.ibm.com/comm/research_projects.nsf/pages/hpct.index.html • Scalable Performance Analysis of Large-Scale Applications (SCALASCA) Toolset: Scalasca is an open-source toolset that can be used to analyze the performance behavior of parallel applications and to identify opportunities for optimization. It has been specifically designed for use on large-scale systems including BlueGene and Cray XT, but is also well-suited for small- and medium-scale HPC platforms. Scalasca supports an incremental performance-analysis procedure that integrates runtime summaries with in-depth studies of concurrent behavior via event tracing, adopting a strategy of successively refined measurement configurations. A distinctive feature is the ability to identify wait states that occur, for example, as a result of unevenly distributed workloads. Especially when trying to scale communication-intensive applications to large processor counts, such wait states can present severe challenges to achieving good performance. Scalasca is developed by the Julich Supercomputing Centre and available under the New BSD open-source license. More information can be found here: http://www.scalasca.org • CEPBA Toolkit: The CEPBA-tools environment is a trace based analysis environment consisting with two major components, Paraver, a browser for traces obtained from a parallel run and Dimemas , a simulator to rebuild the time behavior of a parallel program from a trace. More information can be found here: http://www.bsc.es/plantillaF.php?cat_id=52
I-Hsin Chung, Seetharami R. Seelam, Bernd Mohr, Jesús Labarta
IPDPS1
2009 Towards a framework for automated performance tuning
abstract
As part of the DARPA sponsored high productivity computing systems (HPCS) program, IBM is building petaflop supercomputers that will be fast, power-efficient, and easy to program. In addition to high performance, high productivity to the end user is another prominent goal. The challenge is to develop technologies that bridge the productivity gap - the gap between the hardware complexity and the software limitations. In addition to language, compiler, and runtime research, powerful and user-friendly performance tools are critical in debugging performance problems and tuning for maximum performance. Traditional tools have either focused on specific performance aspects (e.g., communication problems) or provided limited diagnostic capabilities, and using them alone usually do not pinpoint accurately performance problems. Even fewer tools attempt to provide solutions for problems detected. In our study, we develop an open framework that unifies tools, compiler analysis, and expert knowledge to automatically analyze and tune the performance of an application. Preliminary results demonstrated the efficiency of our approach.
Guojing Cong, Seetharami R. Seelam, I-Hsin Chung, Sophia Wen, David J. Klepacki
IPDPS3
2009 Application level I/O caching on Blue Gene/P systems
abstract
In this paper, we present an application level aggressive I/O caching and prefetching system to hide I/O access latency experienced by out-of-core applications. Without the application level prefetching and caching capability, users of I/O intensive applications need to rewrite them with asynchronous I/O calls or restructure their code with MPI-IO calls to efficiently use the large scale system resources. Our proposed solution of user controllable aggressive caching and prefetching system maintains a file-IO cache in the user space of the application, analyzes the I/O access patterns, prefetches requests, and performs write-back of dirty data to storage asynchronously. So each time the application needs the data it does not have to pay the full I/O latency penalty in going to the storage and getting the required data. We have implemented this aggressive caching and asynchronous prefetching on the Blue Gene/P (BGP) system. The preliminary experiment evaluates the caching performance using the WRF benchmark. The results on BGP system demonstrate that our method improves application I/O throughput.
Seetharami R. Seelam, I-Hsin Chung, John Bauer, Hao Yu 0008, Hui-Fang Wen
IPDPS2
2008 Workload Performance Characterization of DARPA HPCS Benchmarks
abstract
It is critical to understand the workload characteristics and resource usage patterns of available applications to guide the design and development of hardware and software stacks of the future machines. In this paper, we analyze the workload performance characteristics of three large-scale DARPA HPCS benchmarks: HYCOM, POP, and LBMHD while executing on IBM Power5+ processor machines. Our analysis is focused on CPU/memory performance using cycles per instruction (CPI) model and multiprocess communication performance using MPI traces. For each benchmark, we provide a high level performance analysis followed by the hot-spot analysis of codes for selected input parameters.Then we present a detailed workload performance characterization using CPI model with data from a unique set of performance counters available on the Power5+ processor system. For communication, we describe the sources of load imbalances in the applications and identify the potential impediments to scalability of the applications under large processor counts.We identify several sources of performance problems that are potential bottlenecks and discuss methods to ameliorate them.
Seetharami R. Seelam, I-Hsin Chung, Guojing Cong, Hui-Fang Wen, David J. Klepacki
HPCC2
2008 A framework for automated performance bottleneck detection
abstract
In this paper, we present the architecture design and implementation of a framework for automated performance bottleneck detection. The framework analyzes the time-spent distribution in the application and discovers the performance bottlenecks by using given bottleneck definitions. The user can query the application execution performance to identify performance problems. The design of the framework is flexible and extensible so it can be tailored based on the actual application execution environment and performance tuning requirement. To demonstrate the usefulness of the framework, we apply the framework on a practical DARPA application and show how it helps to identify performance bottlenecks. The framework helps to automate the performance tuning process and improve the user’s productivity.
I-Hsin Chung, Guojing Cong, David J. Klepacki, Simone Sbaraglia, Seetharami R. Seelam, Hui-Fang Wen
IPDPS1
2008 Early experiences in application level I/O tracing on blue gene systems
abstract
On todays massively parallel processing (MPP) supercomputers, it is increasingly important to understand I/O performance of an application both to guide scalable application development and to tune its performance. These two critical steps are often enabled by performance analysis tools to obtain performance data on thousands of processors in an MPP system. To this end, we present the design, implementation, and early experiences of an application level I/O tracing library and the corresponding tool for analyzing and optimizing I/O performance on Blue Gene (BG) MPP systems. This effort was a part of IBM HPC Toolkit for BG systems. To our knowledge, this is the first comprehensive application-level I/O monitoring, playback, and optimizing tool available on BG systems. The preliminary experiments on popular NPB BTIO benchmark show that the tool is much useful on facilitating detailed I/O performance analysis.
Seetharami R. Seelam, I-Hsin Chung, Ding-Yong Hong, Hui-Fang Wen, Hao Yu 0008
IPDPS2
2007 Performance Studies of a WebSphere Application, Trade, in Scale-out and Scale-up Environments
abstract
Scale-out approach, in contrast to scale-up approach (exploring increasing performance by utilizing more powerful shared-memory servers), refers to deployment of applications on a large number of small, inexpensive, but tightly packaged and tightly interconnected servers. Recently, there has been an increasing interest in scale-out approach. The purpose of this study is to discover advantages or disadvantages of scale-out systems with a typical enterprise workload, IBM Trade Performance Benchmark Sample for Websphere application server (a.k.a. Trade6). In this work, through cross system performance comparison, we show that for such workload, scale-out approach has better performance/cost effect. In term of scalability, we show that Websphere application server packages for distributed environment scale well while the possible bottleneck of the application deployment is the database tier. We present preliminary results to show that both database partitioning feature (DPF) and federated database server approaches are not exactly suitable for providing scale-out solution for the database tier of workloads similar to Trade (small tables and short transactions). In addition, we discuss our on-going effort on further performance study: (1) studies of performance/scalability for larger deployments by adopting the IBM AMBIENCE queuing network modeling tool, (2) performance breakdowns utilizing IBM ACTC hardware counter library.
Hao Yu 0008, José E. Moreira, Parijat Dube, I-Hsin Chung, Li Zhang 0002
IPDPS4
2006 A Case Study Using Automatic Performance Tuning for Large-Scale Scientific Programs
abstract
Active Harmony is an automated runtime performance tuning system. In this paper we describe several case studies of using Active Harmony to improve the performance for scientific libraries and applications. We improved the tuning mechanism so it can work iteratively with benchmarking runs. By tuning the computation and data distribution, Active Harmony helps applications that utilize the PETSc library to achieve better load balance and to reduce the execution time up to 18%. For the climate simulation application POP using 480 processors, the tuning results show that by changing the block size and parameter values, the execution time is reduced up to 16.7%. Active Harmony is able to improve GS2, a plasma physics code, up to a factor of 5.1 times faster. The experiment results show that the Active Harmony system is a feasible and useful tool to automated performance tuning for scientific libraries and applications
I-Hsin Chung, Jeffrey K. Hollingsworth
HPDC1
2006 A study of MPI performance analysis tools on Blue Gene/L
abstract
Applications on today's massively parallel supercomputers rely on performance analysis tools to guide them toward scalable performance on thousands of processors. However, conventional tools for parallel performance analysis have serious problems due to the large data volume that may be required. In this paper, we discuss the scalability issue for MPI performance analysis on Blue Gene/L, the world's fastest supercomputing platform. We present an experimental study of existing MPI performance tools that were ported to BG/L from other platforms. These tools can be classified into two categories: profiling tools that collect timing summaries, and tracing tools that collect a sequence of time-stamped events. Profiling tools produce small data volumes and can scale well, but tracing tools tend to scale poorly. The experimental study discusses the advantages and disadvantages for the tools in the two categories and will be helpful in the future performance tools design.
I-Hsin Chung, Robert Walkup, Hui-Fang Wen, Hao Yu 0008
IPDPS1
2006 MPI tools and performance studies - MPI performance analysis tools on Blue Gene/L
abstract
Applications on today's massively parallel supercomputers are often guided with performance analysis tools toward scalable performance on thousands of processors. However, conventional tools for parallel performance analysis have serious problems due to the large data volume that needs to be handled. In this paper, we discuss the scalability issue for MPI performance analysis on Blue Gene/L, the world's fastest supercomputing platform. First we present an experimental study of existing MPI performance tools that were ported to BG/L from other platforms. These tools can be classified into two categories: profiling tools that collect timing summaries, and tracing tools that collect a sequence of time-stamped events. Profiling tools produce small data volumes and can scale well, but tracing tools tend to scale poorly. We then describe a configurable MPI tracing tool developed for BG/L. By providing a configurable method for trace generation. the volume of trace data can be controlled, and scalability is significantly improved.
I-Hsin Chung, Robert Walkup, Hui-Fang Wen, Hao Yu 0008
SC1
2006 Blue Gene system software - Topology mapping for Blue Gene/L supercomputer
abstract
Mapping virtual processes onto physical processos is one of the most important issues in parallel computing. The problem of mapping of processes/tasks onto processors is equivalent to the graph embedding problem which has been studied extensively. Although many techniques have been proposed for embeddings of two-dimensional grids, hypercubes, etc., there are few efforts on embeddings of three-dimensional grids and tori. Motivated for better support of task mapping for Blue Gene/L supercomputer, in this paper, we present embedding and integration techniques for the embeddings of three-dimensional grids and tori. The topology mapping library that based on such techniques generates high-quality embeddings of two/three-dimensional grids/tori. In addition, the library is used in BG/L MPI library for scalable support of MPI topology functions. With extensive empirical studies on large scale systems against popular benchmarks and real applications, we demonstrate that the library can significantly improve the communication performance and the scalability of applications.
Hao Yu 0008, I-Hsin Chung, José E. Moreira
SC2
2004 Automated Cluster-Based Web Service Performance Tuning
I-Hsin Chung, Jeffrey K. Hollingsworth
HPDC1
2004 Using Information from Prior Runs to Improve Automated Tuning Systems
abstract
Active Harmony is an automated runtime performance tuning system. In this paper we describe a parameter prioritizing tool to help focus on those parameters that are performance critical. Historical data is also utilized to further speed up the tuning process. We first verify our proposed approaches with synthetic data and finally we verify all the improvements on a real cluster-based web service system. Taken together, these changes allow the Active Harmony system to reduce the time spent tuning from 35% up to 50% and at the same time, reduce the variation in performance while tuning.
I-Hsin Chung, Jeffrey K. Hollingsworth
SC1
2002 Active harmony: towards automated performance tuning
abstract
In this paper, we present the Active Harmony automated runtime tuning system. We describe the interface used by programs to make applications tunable. We present the Library Specification Layer which helps program library developers expose multiple variations of the same API using different algorithms.The Library Specification Language helps to select the most appropriate program library to tune the overall performance. We also present the optimization algorithm used to adjust parameters in the application and the libraries. Finally, we present results that show how the system is able to tune several real applications. The automated tuning system is able to tune the application parameers to within a few percent of the best value after evaluating only 11 out of over 1,700 possible configurations.
Cristian Tapus, I-Hsin Chung, Jeffrey K. Hollingsworth
SC2
2002 Design of Scalable Continuous Media Servers
Cheng-Fu Chou, Leana Golubchik, John C. S. Lui, I-Hsin Chung
Multim. Tools Appl.4