VLDB 2026 Research / reviewers in the wild / expert
Jerry Chou 0001
dblp:80/9883 · also Jerry Chi-Yuan Chou
· DBLP profile ↗
64ranked-venue papers
11as first author
27since 2021 · last 2026
0000-0001-7851-1140ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 6 first-author · 11 since 2021Computer networks · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-authorSecurity and privacy · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HybridServe: Stall-Free Distributed Disaggregated LLM Inference with Hybrid KVCache Buffering
Yi-Syuan Ke, Zhan-Wei Wu, Chih-Tai Tsai, Sao-Hsuan Lin, Jerry Chou 0001 |
CLOSER | 5 |
| 2026 | SATORU: Proactive Length-Aware Scheduling for High-Throughput Batch LLM Serving
Gideon Levi, Jerry Chou 0001 |
CLOSER | 2 |
| 2026 | NIKA: Optimal KV Cache Transfer for Minimizing the Latency of Disaggregated LLM Inference
Chih-Tai Tsai, Zhan-Wei Wu, Yi-Syuan Ke, Sao-Hsuan Lin, Jerry Chou 0001 |
CLOSER | 5 |
| 2026 | MoE-Pipe: A Pipelined MoE Model Loading Framework for Reducing the Cold Start Delays in Serverless Inference
Zhan-Wei Wu, Chih-Tai Tsai, Sao-Hsuan Lin, Yi-Syuan Ke, Jerry Chou 0001 |
CLOSER | 5 |
| 2025 | WATCH '25: First Workshop on Analytics, Telemetry, and Cybersecurity for HPCCabstractThe Workshop on Analytics, Telemetry, and Cybersecurity for High-Performance Computing and Communications (HPCC) is newly launched and takes place in Taipei, Taiwan, on October 17, 2025, in conjunction with the ACM Conference on Computer and Communications Security (CCS'25). As its title suggests, the workshop centers on strengthening resilience and security in HPCC applications and infrastructures by leveraging leading technologies, including data-driven methodologies and machine intelligence techniques. The primary objective of this workshop is to provide a dedicated platform for researchers, practitioners, and industry experts to engage in discussions on cutting-edge topics in analytics, telemetry, and cybersecurity for HPCC. This year's call for contributions welcomes both full research papers and work-in-progress submissions, resulting in the acceptance of five full-length research papers. In addition, the workshop features a distinguished keynote presentation by Dr. Hsu-Chun Hsiao, Associate Professor in the Department of Computer Science and Information Engineering and the Graduate Institute of Networking and Multimedia at National Taiwan University. The WATCH'25 complete workshop proceedings can be found at: https://dl.acm.org/citation.cfm?id=3733826. Massimo Cafaro, Eric Chan-Tin, Jerry Chou 0001, Jinoh Kim |
CCS | 3 |
| 2025 | License-Aware Task Clustering and Adaptive DAG Restructuring for EDA Scheduling on Hybrid CloudabstractThe Electronic Design Automation (EDA) industry faces increasing computational demands in the AI era, yet existing scheduling algorithms inadequately address the unique challenges of hybrid cloud EDA workflows, particularly software license constraints and atomic workflow execution requirements. While previous research focuses on task-level scheduling with fine-grained preemption, many production EDA environments orchestrate jobs at the workflow-level, requiring coarse-grained scheduling. As a result, optimizing license usage becomes sig-nificantly harder, since each entire workflow must be sched-uled as an atomic unit. This paper presents a comprehensive framework combining a multi-objective ranking function for workflow prioritization, License-Aware Task Clustering (LATC) with Two-Phase License Early Releasing mechanism to enhance license circulation, and adaptive DAG restructuring that reduces resource requirements by trading parallelism for feasibility in a limited environment. Experimental results demonstrate 65 % reduction in average wait time and decreased deadline miss rates from 62.5% to 12.5% compared to conventional scheduling baseline through our license management strategies. Wei-Sheng Lin, Yu-Chen Yeh, Chao-Hung Chen, Te-Yen Liu, Jerry Chou 0001 |
CloudCom | 5 |
| 2025 | Fast Malicious Packets Inspection Framework Using Converged Accelerator
Chuan-Ming Ou, Yong-Xuan Huang, Ming-Hung Chen, I-Hsin Chung, Jerry Chou 0001 |
HPC Asia | 5 |
| 2025 | PCIe Bandwidth-Aware Scheduling for Multi-Instance GPUs
Yan-Mei Tang, Wei-Fang Sun, Hsu-Tzu Ting, Ming-Hung Chen, I-Hsin Chung, Jerry Chou 0001 |
HPC Asia | 6 |
| 2025 | Libra: A Python-Level Tensor Re-Materialization Strategy for Reducing Deep Learning GPU Memory Usage
Ling-Sung Wang, Sao-Hsuan Lin, Jerry Chou 0001 |
HPC Asia | 3 |
| 2025 | PBHS: A Prediction-Based Scheduler for Hyperparameter Tuning
Hong-Feng Yu, Cheng-Hsun Chang, Jerry Chou 0001 |
HPC Asia | 3 |
| 2025 | Reproducing Performance of Data-Centric Python by SCC Team From National Tsing Hua UniversityabstractAs part of the Student Cluster Competition at the SC22 conference, this work aims to reproduce the performance evaluations of the Data Centric (DaCe) Python framework by leveraging Intel MKL and NVIDIA CUDA interface. The evaluations are conducted on a single CPU-based node, NVIDIA A100 GPUs, and a eight-node cloud supercomputer. Our experimental results successfully reproduce the performance evaluations on our cluster. Additionally, we provide insightful analysis and propose effective methods for achieving higher performance when utilizing DaCe as an acceleration library. Fu-Chiang Chang, En-Ming Huang, Pin-Yi Kuo, Chan-Yu Mou, Hsu-Tzu Ting, Pang-Ning Wu, Jerry Chou 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2024 | CacheFlow: Enhancing Data Flow Efficiency in Serverless Computing by Local Caching
Yi-Syuan Ke, Zhan-Wei Wu, Chih-Tai Tsai, Sao-Hsuan Lin, Jerry Chou 0001 |
CLOSER | 5 |
| 2023 | Optimal Static Bidding Strategy for Running Jobs with Hard Deadline Constraints on Spot Instances
Kai-Siang Wang, Cheng-Han Hsieh, Jerry Chou 0001 |
CLOSER | 3 |
| 2023 | PHY: A performance-driven hybrid communication compression method for distributed training
Chen-Chun Chen, Yu-Min Chou, Jerry Chou 0001 |
J. Parallel Distributed Comput. | 3 |
| 2023 | Gemini: Enabling Multi-Tenant GPU Sharing Based on Kernel Burst EstimationabstractRecent years have seen rapid adoption of GPUs in various types of platforms because of the tremendous throughput powered by massive parallelism. However, as the computing power of GPU continues to grow at a rapid pace, it also becomes harder to utilize these additional resources effectively with the support of GPU sharing. In this work, we designed and implementedGemini, a user-space runtime scheduling framework to enable fine-grained GPU allocation control with support for multi-tenancy and elastic allocation, which are critical for cloud and resource providers. Our key idea is to introduce the concept ofkernel burst, which refers to a group of consecutive kernels launched together without being interrupted by synchronous events. Based on the characteristics of kernel burst, we proposed a low overheadevent-driven monitorand adynamic time-sharing schedulerto achieve our goals. Our experiment evaluations using five types of GPU applications show that Gemini enabled multi-tenant and elastic GPU allocation with less than 5% performance overhead. Furthermore, compared to static scheduling, Gemini achieved 20%$\sim$30% performance improvement without requiring prior knowledge of applications. Hung-Hsin Chen, En-Te Lin, Yu-Min Chou, Jerry Chou 0001 |
IEEE Trans. Cloud Comput. | 4 |
| 2022 | HPA: Hierarchical Placement Algorithm for Multi-Cloud Microservices ApplicationsabstractMicroservices have become a popular way to develop and operate complex (mostly web-based) applications by decomposing software into small loosely coupled services communicating over well-defined APIs. It offers several benefits, including manageability, flexibility and scalability to name a few. However, the orchestration of microservices also becomes more complex and challenging. One of the critical problems is the placement of services which can greatly affect the communication cost and application performance. Especially, with the growing trend of multi-cloud adaptation for enterprise companies, the scale and impact of the placement problem also grows significantly. To tackle the problem, we conduct experiments on public cloud (AWS) to analyze and model the performance impact of placement decision. Then we propose a Hierarchical Placement algorithm (HPA), combining genetic algorithm and spectral clustering technique, to achieve fast and accurate placement decision. Our evaluation results reveal that HPA reduces the service request turn-around time of random placement by up to 44.66%, and only 16% slower than the optimal solution obtained by an ILP solver. HanTing Liang, Jerry Chou 0001 |
CloudCom | 2 |
| 2022 | SNTA'22: The 5th Workshop on Systems and Network Telemetry and AnalyticsabstractHPC and distributed systems are the driving force for the advancement of many emerging technologies, such as exascale systems, quantum machines, terabit networking, 5G/6G wireless, and cloud/edge computing. The tasks of systems and network telemetry are a key element for effective operations and management of the advancement of many emerging systems and technologies, and require more scalable telemetry and analysis techniques for comprehensive monitoring and analysis. Various input sources such as end systems, switches, firewalls, intrusion sensors and the emerging network elements speaking with different syntax and semantics make organizing and incorporating the generated data challenging for the quantitative and qualitative analysis. This workshop looks for new approaches and methods at the intersection of HPC systems and data sciences to address these difficult challenges of emerging technologies from the diverse angles of systems/network performance, availability, reliability, and security. Jinoh Kim, Massimo Cafaro, Jerry Chou 0001, Alex Sim |
HPDC | 3 |
| 2022 | A Reservation-Based List Scheduling for Embedded Systems with Memory Constraints
Kai-Siang Wang, Jerry Chou 0001 |
PDCAT | 2 |
| 2022 | ALBERT: An automatic learning based execution and resource management system for optimizing Hadoop workload in clouds
Chen-Chun Chen, Kai-Siang Wang, Yu-Tung Hsiao, Jerry Chou 0001 |
J. Parallel Distributed Comput. | 4 |
| 2022 | Performance benchmarking and auto-tuning for scientific applications on virtual cluster
Ke-Jou Hsu, Jerry Chou 0001 |
J. Supercomput. | 2 |
| 2022 | Optimization of multi-class 0/1 knapsack problem on GPUs by improving memory access efficiency
En-Ming Huang, Jerry Chou 0001 |
J. Supercomput. | 2 |
| 2022 | CAMIRA: a consolidation-aware migration avoidance job scheduling strategy for virtualized parallel computing clusters
Satyajit Padhy, Ming-Han Tsai, Jerry Chou 0001 |
J. Supercomput. | 4 |
| 2021 | DynamoML: Dynamic Resource Management Operators for Machine Learning Workloads
Min-Chi Chiang, Jerry Chou 0001 |
CLOSER | 2 |
| 2021 | A Deep Reinforcement Learning Method for Solving Task Mapping Problems with Dynamic Traffic on Parallel SystemsabstractEfficient mapping of application communication patterns to the network topology is a critical problem for optimizing the performance of communication bound applications on parallel computing systems. The problem has been extensively studied in the past, but they mostly formulate the problem as finding an isomorphic mapping between two static graphs with edges annotated by traffic volume and network bandwidth. But in practice, the network performance is difficult to be accurately estimated, and communication patterns are often changing over time and not easily obtained. Therefore, this work proposes a deep reinforcement learning (DRL) approach to explore better task mappings by utilizing the performance prediction and runtime communication behaviors provided from a simulator to learn an efficient task mapping algorithm. We extensively evaluated our approach using both synthetic and real applications with varied communication patterns on Torus and Dragonfly networks. Compared with several existing approaches from literature and software library, our proposed approach found task mappings that consistently achieved comparable or better application performance. Especially for a real application, the average improvement of our approach on Torus and Dragonfly networks are 11% and 16%, respectively. In comparison, the average improvements of other approaches are all less than 6%. YuCheng Wang, Jerry Chou 0001, I-Hsin Chung |
HPC Asia | 2 |
| 2021 | MIRAGE: A consolidation aware migration avoidance genetic job scheduling algorithm for virtualized data centers
Satyajit Padhy, Jerry Chou 0001 |
J. Parallel Distributed Comput. | 2 |
| 2021 | Distributed and incremental travelling salesman algorithm on time-evolving graphs
Jerry Chou 0001 |
J. Supercomput. | 2 |
| 2021 | Critique of "Planetary Normal Mode Computation: Parallel Algorithms, Performance, and Reproducibility" by SCC Team From National Tsing Hua UniversityabstractAs a special activity of the Student Cluster Competition at SC19 conference, we made an attempt to reproduce the scalability evaluations of a highly paralleled polynomial filtering eigensolver for computing planetary interior normal modes. Our experiments were conducted on a Mars dataset using a small scale 4-node cluster with Intel Skylake CPU architecture, while the original article's were conducted on a Moon dataset using a large scale 256-node supercomputer with Intel CPU Skylake and KNL architectures. This article shares our experiences and observations from our reproducibility activity and discusses our findings on three main sections: the weak scalability, the strong scalability, and the relationships between variables. The results of weak scalability and strong scalability were successfully reproduced. But due to the differences on the problem scale, input dataset, and system architecture, different behaviors regarding the polynomial degree were observed. Wei-Fang Sun, Hung-Hsin Chen, ShaoFu Lin, YuanChing Lin, Jing-Wei Wu, En-Te Lin, Jerry Chou 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2020 | ECS2: A Fast Erasure Coding Library for GPU-Accelerated Storage Systems with Parallel & Direct IOabstractAs data volume keeps increasing at a rapid rate, there is an urgent need for large, reliable, and cost-effective storage systems. Erasure coding has drawn increasing attention because of its ability to ensure data reliability with higher storage efficiency, and it has been widely adopted in many distributed and large-scale storage systems, such as Azure cloud storage and HDFS. However, the storage efficiency of erasure code comes at the price of higher computing complexity. While many studies have shown the coding computations can be significantly accelerated using GPU, the overhead of data transfer between storage devices and GPUs become a new performance bottleneck. In this work, we designed and implemented, ECS2, a fast erasure coding library on GPU -accelerated storage to let users enhance their data protection with transparent IO performance and file system like programming interface. By taking advantage of the latest GPUDirect technology supported on Nvidia GPU, our library is able to bypass CPU and host memory copy from the IO path, so that both the computing and IO overhead from coding can be minimized. Using synthetic IO workload based on real storage system trace, we show that the IO latency can be reduced by 10% ~ 20% with GPUDirect technology, and the overall IO throughput of a storage system can be improved up to 70%. ChanJung Chang, Jerry Chou 0001, Yu-Ching Chou, I-Hsin Chung |
CLUSTER | 2 |
| 2020 | KubeShare: A Framework to Manage GPUs as First-Class and Shared Resources in Container CloudabstractContainer has emerged as a new technology in clouds to replace virtual machines~(VM) for distributed applications deployment and operation. With the increasing number of new cloud-focused applications, such as deep learning and high performance applications, started to reply on the high computing throughput of GPUs, efficiently supporting GPU in container cloud becomes essential. While GPU virtualization has been extensively studied for VM, limited work has been done for containers. One of the key challenges is the lack of support for GPU sharing between multiple concurrent containers. This limitation leads to low resource utilization when a GPU device cannot be fully utilized by a single application due to the burstiness of GPU workload and the limited memory bandwidth. To overcome this issue, we designed and implemented KubeShare, which extends Kubernetes to enable GPU sharing with fine-grained allocation. KubeShare is the first solution for Kubernetes to make GPU device as a first class resources for scheduling and allocations. Using real deep learning workloads, we demonstrated KubeShare can significantly increase GPU utilization and overall system throughput around 2x with less than 10% performance overhead during container initialization and execution. Ting-An Yeh, Hung-Hsin Chen, Jerry Chou 0001 |
HPDC | 3 |
| 2019 | DRAGON: A Dynamic Scheduling and Scaling Controller for Managing Distributed Deep Learning Jobs in Kubernetes Cluster
Chan-Yi Lin, Ting-An Yeh, Jerry Chou 0001 |
CLOSER | 3 |
| 2019 | Student Cluster Competition 2018, team NTHU: Reproducing performance of multi-physics simulations of the tsunamigenic 2004 sumatra megathrust earthquake on the Intel Skylake architecture
ShaoFu Lin, ChiChen Yang, Scott Cheng, KengJui Hsu, Hung-Hsin Chen, YuanChing Lin, Jerry Chou 0001 |
Parallel Comput. | 7 |
| 2018 | Comparison Between Bare-metal, Container and VM using Tensorflow Image Classification Benchmarks for Deep Learning Cloud Platform
Chan-Yi Lin, Hsin-Yu Pai, Jerry Chou 0001 |
CLOSER | 3 |
| 2018 | A Computation Workload Characteristic Study of C-RANabstractDriven by the surging demand of mobile applications and IoT devices, the amount of global mobile data traffic is estimated to increase sevenfold in the next few years and reach 69 exabytes per month by 2022. This rapid growing rate force telecom operator to adapt new wireless network technologies in order to deliver desired network performance and quality while reduce network deployment and operating costs. One of the approaches that has gained more traction recently is C-RAN, which aims to renovate the infrastructure of radio access network based on cloud technology. In this work, we built a cloudified LTE testbed environment of C-RAN by integrating the OpenAirInterface (OAI), an open-source software radio solution, with the OpenStack, and open-source cloud infrastructure solution. Using the testbed, we conducted workload study to understand the computation resource demand of C-RAN software, and proposed a function splitting technique to improve the resource utilization of C-RAN cloud platform. Yu-Cing Luo, Shih-Chun Huang, Jerry Chou 0001, Bing-Liang Chen |
ICDCS | 3 |
| 2018 | Student cluster competition 2017, team NTHU: Reproducing vectorization of the tersoff multi-body potential on the Intel Skylake and Nvidia P100 architecture
ChanJung Chang, YungChing Lin, Scott Cheng, YuCheng Wang, LiYu Yu, TienChi Yang, Jerry Chou 0001 |
Parallel Comput. | 7 |
| 2017 | Apply Block Index Technique to Scientific Data Analysis and I/O SystemsabstractScientific discoveries are increasingly relying on analysis of massive amounts of data. The ability to directly access the most relevant data records through query, without shifting through all of them becomes essential. However, scientific datasets are commonly stored on parallel file systems and I/O systems that are optimized for reading/writing large chunks of data, and many scientific datasets have spatial-temporal data similarity, such that the records with similar values often locate in a close proximity of each other. Therefore, our previous work started to investigate the benefit of using block range index technique for scientific datasets, which only records the value range of all the records in a data block. In this paper, we extend our work in several aspects. First, we implement and integrate our blockindex technique with the ADIOS I/O system. Second, we show our proposed method can be significantly better than the existing minmax and bitmaps indexing methods supported in ADIOS, and can also have comparable performance in the worst case. Third, we propose several techniques that can take advantage of the block index information to greatly reduce data retrieval time from query results. Fourth, we evaluate our approach using several real scientific datasets, and analyze the spatial-temporal data similarity characteristics in them. Through our study, we believe block index can be an effective indexing technique for scientific datasets with little implementation and operating overhead. It's size is small enough for building the indexes on-the-fly, and yet its query information is sufficient for efficient data access. Tzu-Hsien Wu, Jerry Chou 0001, Norbert Podhorszki, Junmin Gu, Yuan Tian 0004, Scott Klasky, Kesheng Wu |
CCGrid | 2 |
| 2017 | Strike the Balance between System Utilization and Data Locality under Deadline Constraint for MapReduce ClustersabstractMapReduce paradigm has become a popular platform for massive data processing and Big Data applications. Although MapReduce was initially designed for high throughput and batch processing, it has also been used for handling many other types of applications and workloads due to its scalable and reliable system architecture. One of the emerging requirements for enterprise data-process computing is completion time guar- antee. However, there are only a few research works have been done for MapReduce jobs with deadline constraint. Therefore, in this paper, we aim to prevent jobs from missing deadline while maximizing the resource utilization and data locality of a MapReduce cluster. Our approach is to introduce a two-phase job scheduling mechanism which combines a job admission controller policy and a priority-based scheduling algorithm. We use a series of simulations over diverted workload to evaluate our system. The results show that our approach can guarantee job completion time in a heavy-loaded system, and achieve comparable data locality to the delay schedule algorithm in a light-loaded system. Furthermore, our approach can maximize system throughput by preventing system resources from being wasted by the jobs missing their deadlines. Yeh-Cheng Chen, Jerry Chou 0001 |
PDCAT | 2 |
| 2017 | Optimizing the query performance of block index through data analysis and I/O modelingabstractIndexing technique has become an efficient tool to enable scientists to directly access the most relevant data records. But, the time and space requirements of building and storing indexes are expensive in the traditional approaches, such as R-tree and bitmaps. Recently, we started to address this issue by using the idea of "block index", and our previous work has shown promising results from comparing it against other well-known solutions, including ADIOS, SciDB, and FastBit. In this work, we further improve the technique from both theoretical and implementation perspectives. Driven by an extensive effort in characterizing scientific datasets and modeling I/O systems, we presented a theoretical model to analyze its query performance with respect to a given block size configuration. We also introduced three optimization techniques to achieve a 2.3x query time reduction comparing to the original implementation. Tzu-Hsien Wu, Jerry Chou 0001, Shyng Hao, Bin Dong 0002, Scott Klasky, Kesheng Wu |
SC | 2 |
| 2016 | Indexing Blocks to Reduce Space and Time Requirements for Searching Large Data FilesabstractScientific discoveries are increasingly relying on analysis of massive amounts of data generated from scientific experiments, observations, and simulations. The ability to directly access the most relevant data records, without shifting through all of them becomes essential. While many indexing techniques have been developed to quickly locate the selected data records, the time and space required for building and storing these indexes are often too expensive to meet the demands of in situ or real-time data analysis. Existing indexing methods generally capture information about each individual data record, however, when reading a data record, the I/O system typically has to access a block or a page of data. In this work, we postulate that indexing blocks instead of individual data records could significantly reduce index size and index building time without increasing the I/O time for accessing the selected data records. Our experiments using multiple real datasets on a supercomputer show that block index can reduce query time by a factor of 2 to 50 over other existing methods, including SciDB and FastQuery. But the size of block index is almost negligible comparing to the data size, and the time of building index can reach the peak I/O speed. Tzu-Hsien Wu, Shyng Hao, Jerry Chou 0001, Bin Dong 0002, Kesheng Wu |
CCGrid | 3 |
| 2016 | Dynamic Block Partitioning Strategy for Cloud-Backed File SystemsabstractCloud storage services, like S3, has drawn increasing popularity as primary data storage repository for users and enterprises because data can be conveniently accessed and shared through simplified web service based interface. However, users still desire the full use of the performance and capacity of their purchased storage resources, and to retain the POSIX I/O interface for accommodating general applications on local systems. Therefore, there are increasing research interests in cloud-backed file systems that aim to achieve the best of both worlds by connecting POSIX-interface local file system with a remote cloud storage system. However, due to the differences between the two types of storage systems, many challenges arise. One of them is to support efficient random file access on the object cloud storage system. Several frameworks have proposed to partition a file into fixed-size blocks and store them as individual objects on cloud for partial data access. But the setting of block size is not trivial and can have significant performance impact. Larger block size may contain data that is not requested by users. Smaller block may require more data access requests and lower network bandwidth utilization. To overcome the demerit of a static fixed-size strategies, this paper proposes a dynamically block partitioning strategy that will merge and split blocks on-the-fly according to the I/O access pattern of a file. Our simulation results show that our strategy can improve I/O throughput by 37% to 321% comparing to the static strategy using different fixed-size settings. Lung-Hsiang Chung, Ching-Feng Lee, Jerry Chou 0001 |
CloudCom | 3 |
| 2016 | DRASH: A Data Replication-Aware Scheduler in Geo-Distributed Data CentersabstractDriven by the trends of BigData and Cloud computing, there is a growing demand for processing and analyzing data that are generated and stored across geo-distributed data centers. However, due to the limited network bandwidth between data centers and the growing data volume spread across different locations, it has become increasingly inefficient to aggregate data and to perform computations at a single data center. An approach that has been commonly used by data-intensive cluster computation systems, like Hadoop, is to distribute computations based on data locality so that data can be processed locally to reduce the network overhead and improve performance. But limited work has been done to adapt and evaluate such technique for geo-distributed data centers. In this paper, we proposed DRASH (Data-Replication Aware Scheduler), a job scheduling algorithm that enforces data locality to prevent data transfer, and exploits data replications to improve overall system performance. Our evaluation using simulations with realistic workload traces shows that DRASH can outperform other existing approaches by 16% to 60% in average job completion time, and achieve greater improvements under higher data replication factors. Moïse W. Convolbo, Jerry Chou 0001, Shihyu Lu, Yeh-Ching Chung |
CloudCom | 2 |
| 2016 | Performance Evaluations of Cloud Radio Access Networks
Mu-Han Huang, Yu-Cing Luo, Chen-Nien Mao, Bing-Liang Chen, Shih-Chun Huang, Jerry Chou 0001, Shun-Ren Yang, Yeh-Ching Chung, Cheng-Hsin Hsu |
QSHINE | 6 |
| 2016 | A Middleware Solution for Optimal Sensor Management of IoT Applications on LTE Devices
Satyajit Padhy, Hsin-Yu Chang, Ting-Fang Hou, Jerry Chou 0001, Chung-Ta King, Cheng-Hsin Hsu |
QSHINE | 4 |
| 2016 | Cost-aware DAG scheduling algorithms for minimizing execution cost on cloud resources
Moïse W. Convolbo, Jerry Chou 0001 |
J. Supercomput. | 2 |
| 2015 | In-memory Query System for Scientific DataseisabstractThe growing gap between compute performance and I/O bandwidth coupled with the increasing data volumes has resulted in a bottleneck to the traditional post-simulation data processing method. Hence in-situ computing and query-driven data analysis are important techniques to minimize data movement. By taking advantage of the growing memory capacity on supercomputers, we developed an in-memory query system for scientific data analysis. Our approach is a combination of bitmap indexing, spatial data layout re-organization, distributed shared memory, and location-aware parallel execution. Our evaluations using real scientific datasets showed that we can aggregate the memory capacity from thousands of computes nodes to analyze a 750GB simulation dataset without transferring data to remote nodes or storage systems. Comparing to traditional solutions based on out-of-core parallel file systems, we achieve significant higher query performance. Hsuan-Te Chiu, Jerry Chou 0001, Venkat Vishwanath, Kesheng Wu |
ICPADS | 2 |
| 2015 | Exploiting Replication for Energy-Aware Scheduling in Disk Storage SystemsabstractThis paper deals with the problem of scheduling requests on disks for minimizing energy consumption. We first analyze several versions of the energy-aware disk scheduling problem based on assumptions on the arrival pattern of the requests. We show that the corresponding optimization problems are NP-complete. Then both optimal and heuristic scheduling algorithms are proposed to maximize the energy saving of a storage system. We evaluate our approach using multiple realistic I/O traces, disk simulator and energy model. The results show that we significantly reduce energy consumption up to 55 percent and achieve fewer disk spin-up/ down operations and shorter request response time as compared to other approaches. Since our approach attempts to dynamically assign each request to an energy optimized location, it can also benefit from other traditional static or semi-static solutions that rely on data placement or migration. Finally, we show that a write offloading technique can also be adapted into our solution to minimize the impact from write re quests in terms of energy consumption and request response time. Jerry Chou 0001, Ting-Hsuan Lai, Jinoh Kim, Doron Rotem |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | Taiwan UniCloud: A Cloud Testbed with Collaborative Cloud ServicesabstractThis paper introduces a prototype of Taiwan UniCloud, a community-driven hybrid cloud platform for academics in Taiwan. The goal is to leverage resources in multiple clouds among different organizations. Each self-managing cloud can join the UniCloud platform to share its resources and simultaneously benefit from other clouds with scale-out capabilities. Accordingly, resources are elastic and sharable with each other such as to afford unexpected resource demands to each cloud. The proposed platform provides a web portal to operate each cloud via a uniform user interface. The construction of virtual clusters with multi-core VMs is supplied for parallel and distributed processing models. An object-based storage system is also delivered to federate different storage providers. This paper not only presents the architectural design of Taiwan UniCloud, but also evaluates the performance to demonstrate the possibility of current implementation. Experimental results show the feasibility of the proposed platform as well as the benefit from the cloud federation. Wu-Chun Chung, Po-Chi Shih, Kuan-Chou Lai, Kuanching Li, Che-Rung Lee, Jerry Chou 0001, Ching-Hsien Hsu, Yeh-Ching Chung |
IC2E | 6 |
| 2014 | Simplifying index file structure to improve I/O performance of parallel indexingabstractComplex indexing techniques are needed to reduce the time of analyzing massive scientific datasets, but generating these indexing data structures can be very time consuming. In this work, we propose a set of strategies to simplify the index file structure and to improve the I/O performance during index construction using FastQuery, which is a parallel indexing and querying system for scientific data. FastQuery has been used to analyze data from various scientific applications, including a trillion plasma particles simulation. To accelerate query process, FastQuery uses FastBit to build indexes, and then stores the indexes into file system through parallel scientific data format libraries, such as HDF5. Although these data format libraries are designed to support more complex multi-dimensional arrays, we observed that it still takes considerable work to map the indexing data structures into arrays, especially on parallel machines. To address this problem, in this paper, we attempt to minimize the I/O time by storing indexes into our self-defined binary data format. By fully controlling the data structure, we can minimize the I/O synchronization overhead and explore more efficient I/O strategy for storing indexes. Our experiments of indexing a trillion particle dataset using 20,000 cores of a supercomputer show that the proposed binary I/O driver can reach 85% of the peak I/O bandwidth on the system, and achieves a speedup of up to 4X in terms of the total execution time comparing to the previous FastQuery implementation with HDF5 I/O driver. Hsuan-Te Chiu, Jerry Chou 0001, Venkat Vishwanath, Surendra Byna, Kesheng Wu |
ICPADS | 2 |
| 2014 | iPACS: Power-aware covering sets for energy proportionality and performance in data parallel computing clusters
Jinoh Kim, Jerry Chou 0001, Doron Rotem |
J. Parallel Distributed Comput. | 2 |
| 2013 | Prevent VM Migration in Virtualized Clusters via Deadline Driven Placement PolicyabstractVM consolidation has been shown as a promising technique for saving energy costs of a data center. It relies on VM migration to move user applications or jobs onto fewer numbers of physical servers during off peak hour. However, VM migration is a costly operation that could cause several concerns, such as performance degradation and system instability. Most existing works were proposed to minimize the migration cost for dynamic consolidation which migrates VM at the runtime when SLA violation or resource under-utilization is detected. In contrast, this paper aims to proactively prevent VM migration for semi-static VM consolidation by proposing a deadline driven VM placement strategy based on the awareness of the server turn-off time and job execution time. We evaluate our approach using a real HPC cluster trace as well as a set of synthetic generated workloads. The results show our approach can significantly reduce the number of migrations by 70% on the real trace. We also demonstrate that our approach can be resilient to different workload patterns by achieving consistent improvement around 50% over all the synthetic workloads. Ming-Han Tsai, Jerry Chou 0001, Jye Chen |
CloudCom (1) | 2 |
| 2013 | Optimizing fastquery performance on lustre file systemabstractFastQuery is a parallel indexing and querying system we developed for accelerating analysis and visualization of scientific data. We have applied it to a wide variety of HPC applications and demonstrated its capability and scalability using a petascale trillion-particle simulation in our previous work. Yet, through our experience, we found that performance of reading and writing data with FastQuery, like many other HPC applications, could be significantly affected by various tunable parameters throughout the parallel I/O stack. In this paper, we describe our success in tuning the performance of FastQuery on a Lustre parallel file system. We study and analyze the impact of parameters and tunable settings at file system, MPI-IO library, and HDF5 library levels of the I/O stack. We demonstrate that a combined optimization strategy is able to improve performance and I/O bandwidth of FastQuery significantly. In our tests with a trillion-particle dataset, the time to index the dataset reduced by more than one half. Kuan-Wu Lin, Surendra Byna, Jerry Chou 0001, Kesheng Wu |
SSDBM | 3 |
| 2012 | Value-based tiering management on heterogeneous block-level storage systemabstractAs the scale of datacenter continues to grow, it is hard to keep servers homogenous, with the same hardware and performance characteristics. Today's datacenters commonly operates on several generations of servers from multiple vendors, and mix both high-end and low-end devices together to deliver service quality requirement with lowest cost. However, the heterogenous environment also complicates the management of the datacenters, especially in terms of resource allocation. In this paper, we focus on the resource allocation of a tightly unified block-level storage with SSD and HDD. We conduct experiments to quantify the performance of difference access patterns on each type storage devices. Then formulate our resource allocation problem into a ILP (Integer Linear Programming), and proposed data migration algorithms based on the observations. We evaluate our solution by implementing a heterogenous storage consist of HDD, SDD and iSCSI HDD, and show the data access response time can be reduced by 27%. Chai-Hao Tsai, Jerry Chou 0001, Yeh-Ching Chung |
CloudCom | 2 |
| 2012 | Parallel I/O, analysis, and visualization of a trillion particle simulationabstractPetascale plasma physics simulations have recently entered the regime of simulating trillions of particles. These unprecedented simulations generate massive amounts of data, posing significant challenges in storage, analysis, and visualization. In this paper, we present parallel I/O, analysis, and visualization results from a VPIC trillion particle simulation running on 120,000 cores, which produces ~30TB of data for a single timestep. We demonstrate the successful application of H5Part, a particle data extension of parallel HDF5, for writing the dataset at a significant fraction of system peak I/O rates. To enable efficient analysis, we develop hybrid parallel FastQuery to index and query data using multi-core CPUs on distributed memory hardware. We show good scalability results for the FastQuery implementation using up to 10,000 cores. Finally, we apply this indexing/query-driven approach to facilitate the first-ever analysis and visualization of the trillion particle dataset. Surendra Byna, Jerry Chou 0001, Oliver Rübel, Prabhat, Homa Karimabadi, William S. Daughton, Vadim Roytershteyn, E. Wes Bethel, Mark Howison, Ke-Jou Hsu, Kuan-Wu Lin, Arie Shoshani, Andrew Uselton, Kesheng Wu |
SC | 2 |
| 2011 | FastQuery: A Parallel Indexing System for Scientific DataabstractModern scientific datasets present numerous data management and analysis challenges. State-of-the-art index and query technologies such as FastBit can significantly improve accesses to these datasets by augmenting the user data with indexes and other secondary information. However, a challenge is that the indexes assume the relational data model but the scientific data generally follows the array data model. To match the two data models, we design a generic mapping mechanism and implement an efficient input and output interface for reading and writing the data and their corresponding indexes. To take advantage of the emerging many-core architectures, we also develop a parallel strategy for indexing using threading technology. This approach complements our on-going MPI-based parallelization efforts. We demonstrate the flexibility of our software by applying it to two of the most commonly used scientific data formats, HDF5 and NetCDF. We present two case studies using data from a particle accelerator model and a global climate model. We also conducted a detailed performance study using these scientific datasets. The results show that FastQuery speeds up the query time by a factor of 2.5x to 50x, and it reduces the indexing time by a factor of 16 on 24 cores. Jerry Chou 0001, Kesheng Wu, Prabhat |
CLUSTER | 1 |
| 2011 | Energy-Aware Scheduling in Disk Storage SystemsabstractThis paper deals with the problem of scheduling requests on disks for minimizing energy consumption. We first analyze several versions of the energy-aware disk scheduling problem based on assumptions on the arrival pattern of the requests. We show that the corresponding optimization problems are NP-complete by reduction to the set cover or the independent set problem. Then both optimal and heuristic scheduling algorithms are proposed to maximize the energy saving of a storage system. Our evaluation results using two realistic traces show that our approach significantly reduces energy consumption up to 55% and achieves fewer disk spin-up/down operations and shorter request response time as compared to other approaches. Jerry Chou 0001, Jinoh Kim, Doron Rotem |
ICDCS | 1 |
| 2011 | Parallel index and query for large scale data analysisabstractModern scientific datasets present numerous data management and analysis challenges. State-of-the-art index and query technologies are critical for facilitating interactive exploration of large datasets, but numerous challenges remain in terms of designing a system for processing general scientific datasets. The system needs to be able to run on distributed multi-core platforms, efficiently utilize underlying I/O infrastructure, and scale to massive datasets. Jerry Chou 0001, Mark Howison, Brian Austin, Kesheng Wu, Ji Qiang, E. Wes Bethel, Arie Shoshani, Oliver Rübel, Prabhat, Robert D. Ryne |
SC | 1 |
| 2011 | FastQuery: A General Indexing and Querying System for Scientific Data
Jerry Chou 0001, Kesheng Wu, Prabhat |
SSDBM | 1 |
| 2011 | Energy Proportionality and Performance in Data Parallel Computing Clusters
Jinoh Kim, Jerry Chou 0001, Doron Rotem |
SSDBM | 2 |
| 2010 | Birkhoff-von Neumann switching with statistical traffic profiles
Jerry Chou 0001, Bill Lin 0001 |
Comput. Commun. | 1 |
| 2009 | Optimal multi-path routing and bandwidth allocation under utility max-min fairnessabstractAn important goal of bandwidth allocation is to maximize the utilization of network resources while sharing the resources in a fair manner among network flows. To strike a balance between fairness and throughput, a widely studied criterion in the network community is the notion of max-min fairness. However, the majority of work on max-min fairness has been limited to the case where the routing of flows has already been defined and this routing is usually based on a single fixed routing path for each flow. In this paper, we consider the more general problem in which the routing of flows, possibly over multiple paths per flow, is an optimization parameter in the bandwidth allocation problem. Our goal is to determine a routing assignment for each flow so that the bandwidth allocation achieves optimal utility max-min fairness with respect to all feasible routings of flows. We present evaluations of our proposed multi-path utility max-min fair allocation algorithms on a statistical traffic engineering application to show that significantly higher minimum utility can be achieved when multi-path routing is considered simultaneously with bandwidth allocation under utility max-min fairness, and this higher minimum utility corresponds to significant application performance improvements. Jerry Chou 0001, Bill Lin 0001 |
IWQoS | 1 |
| 2009 | Proactive surge protection: a defense mechanism for bandwidth-based attacks
Jerry Chou 0001, Bill Lin 0001, Subhabrata Sen, Oliver Spatscheck |
IEEE/ACM Trans. Netw. | 1 |
| 2008 | Proactive Surge Protection: A Defense Mechanism for Bandwidth-Based Attacks
Jerry Chou 0001, Bill Lin 0001, Subhabrata Sen, Oliver Spatscheck |
USENIX Security Symposium | 1 |
| 2006 | SCALLOP: A Scalable and Load-Balanced Peer-to-Peer Lookup ProtocolabstractA number of structured peer-to-peer (P2P) lookup protocols have been proposed recently. A P2P lookup protocol routes a lookup request to its target node in a P2P distributed system. Existing protocols achieve balanced routing traffic among nodes by assuming that lookup requests are evenly targeted at every node. However, when lookup requests concentrate on a few nodes simultaneously, these nodes become hot spots. Due to uneven routing patterns in existing protocols, hot spots cause unbalanced routing traffic which leads to routing bottlenecks. In this paper, we present a novel structured P2P lookup protocol called SCALLOP that delivers balanced routing and avoids routing bottlenecks at occurrences of hot spots. Among existing protocols, SCALLOP is the first one to accomplish this goal at the fundamental nature of a routing protocol. SCALLOP achieves balanced routing by uniquely constructing a balanced lookup tree for each node. The balanced tree evenly distributes routing traffic among sibling nodes and, therefore, avoids or reduces routing bottlenecks. In addition, as a load-balanced protocol, SCALLOP delivers asymptotically optimal lookup performance at the tradeoff between routing path and routing table size. We conducted a set of simulations to demonstrate the effectiveness of SCALLOP. The results show that, compared-with a most-referenced and representative structured P2P lookup, protocol and a graph-based extension of this protocol, SCALLOP significantly reduces routing bottlenecks while all three protocols deliver comparable lookup performance. Jerry Chou 0001, Kuang-Li Huang, Tsung-Yen Chen |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2004 | SCALLOP: a scalable and load-balanced peer-to-peer lookup protocol for high-performance distributed systemsabstractMany large-scale servers are implemented in a peer-to-peer distributed system. Achieving rapid response time in such a system relies on a scalable lookup protocol to efficiently locate the requested items and a load-balancing mechanism to avoid the hot spot problem by evenly distributing lookup requests. We present a peer-to-peer lookup protocol that addresses both issues. Our lookup protocol uses the technique of distributed hash table (DHT) to store in each node O(logN) routing information in an N-node distributed system. Unlike other DHT-based peer-to-peer protocols, our protocol constructs a balanced lookup tree to avoid hot spots. We compare our protocol with a commonly used peer-to-peer lookup protocol. The experimental results show that our protocol reduces up to 51% of the lookup requests on heavily loaded nodes. Furthermore, our protocol reduces the total lookup requests and thus delivers better performance. Jerry Chou 0001, Kuang-Li Huang |
CCGRID | 1 |
| 2004 | LessLog: A Logless File Replication Algorithm for Peer-to-Peer Distributed SystemsabstractSummary form only given. The technique of replicating frequently-accessed files to other nodes has been widely used in a high-performance distributed system to reduce the load of the nodes hosting these files. Traditional file replication algorithms rely on the analysis of client-access logs to determine the location of the replicated nodes. We present LessLog, a loglessfile replication algorithm, developed for a peer-to-peer distributed system. We first construct a lookup tree for each node. LessLog uses bitwise operations to determine the location of the replicated node without any client-access history. In addition, each replication is guaranteed to reduce the workload of the replicating node by half. A fault-tolerant LessLog model is also presented. The experimental results show that LessLog successfully and efficiently reduces the load of overloaded nodes. Kuang-Li Huang, Jerry Chou 0001 |
IPDPS | 3 |