VLDB 2026 Research / reviewers in the wild / expert
Che-Rung Lee
dblp:64/1249
· DBLP profile ↗
49ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0003-3940-4478ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 10 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 4 since 2021Software engineering, systems software and programming languages · 5 · 3 since 2021Computer networks · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DATT: Dimension-Augmented Tensor-Train Decomposition for Neural Network Compression
Yu-Chuan Tai, Cheng-Yu Sie, Che-Rung Lee |
COMPSAC | 3 |
| 2026 | FEDA: Fast Exponential Decay Approach for Multivariate Time Series Forecasting
Ming-Ting Zhong, Cheng-Yu Sie, Che-Rung Lee |
COMPSAC | 3 |
| 2026 | Data-Driven Lipschitz Continuity: A Cost-Effective Approach to Improve Adversarial RobustnessabstractAs deep neural networks (DNNs) are increasingly deployed in sensitive applications, ensuring their security and robustness has become critical. A major threat to DNNs arises from adversarial attacks, where small input perturbations can lead to incorrect predictions. Recent advances in adversarial training improve robustness by incorporating additional examples from external datasets or generative models. However, these methods often incur high computational costs, limiting their practicality and hindering real-world deployment. In this paper, we propose a cost-efficient alternative based on Lipschitz continuity that achieves robustness comparable to models trained with extensive supplementary data. Unlike conventional adversarial training, our method requires only a single pass over the dataset without gradient estimation, making it highly efficient. Furthermore, our method can integrate seamlessly with existing adversarial training frameworks and enhances the robustness of models without requiring extra generative data. Experimental results show that our approach not only reduces computational overhead but also maintains or improves the defensive capabilities of robust neural networks. This work opens a promising direction for developing practical, scalable defenses against adversarial attacks. Erh-Chung Chen, I-Hsin Chung, Che-Rung Lee |
WACV | 4 |
| 2026 | LTD: Low Temperature Distillation for Gradient Masking-free Adversarial TrainingabstractAdversarial training is a widely adopted strategy to bolster the robustness of neural network models against adversarial attacks. This article revisits the fundamental assumptions underlying image classification and suggests that representing data as one-hot labels is a key factor that leads to vulnerabilities. However, in real-world datasets, data ambiguity often arises, with samples exhibiting characteristics of multiple classes, rendering one-hot label representations imprecise. To address this, we introduce a novel approach, Low-Temperature Distillation (LTD), designed to refine label representations. Unlike previous approaches, LTD incorporates a relatively low temperature in the teacher model, while maintaining a fixed temperature for the student model during both training and inference. This strategy not only refines assumptions about data distribution but also strengthens model robustness and avoids the gradient masking problem commonly encountered in defensive distillation. Experimental results demonstrate the efficacy of the proposed method when combined with existing frameworks, achieving robust accuracy rates of 58.19%, 31.13%, and 42.08% on the CIFAR-10, CIFAR-100, and ImageNet datasets, respectively, without the need for additional data. Erh-Chung Chen, Che-Rung Lee |
ACM Trans. Cyber Phys. Syst. | 2 |
| 2025 | ITTPD: In-place Tensor Transposition with Permutation Decomposition on GPUsabstractTensor transposition is a fundamental operation in tensor computations with broad applications.Most existing algorithms rely on out-of-place transposition, which duplicates the tensor in memory and repositions elements in their transposed order, resulting in high memory demands.On memory-constrained devices such as Graphics Processing Units (GPUs), this approach is often impractical.To address this challenge, we propose ITTPD (In-place Tensor Transposition with Permutation Decomposition), a GPU-based method designed to operate within strict memory constraints.Similarly to the previous work EITHOT, ITTPD utilizes permutation decomposition, breaking down complex permutations into simpler transposition primitives, and determining an execution sequence that fits within available memory.Comparing to existing work, ITTPD partitions the tensor into smaller sections for the tensor size larger than the memory limits and transposes each independently.The implementation is optimized for GPU memory access patterns, improving the execution efficiency for each primitive.Experimental results show that ITTPD significantly outperforms state-of-the-art out-ofplace GPU implementations, offering reduced memory overhead and faster execution times.Moreover, ITTPD can process tensors nearly twice as large as those feasible with out-of-place methods, making it a versatile solution for N-order tensor transpositions. Kai-Jung Cheng, Che-Rung Lee |
HPC Asia | 2 |
| 2024 | Latency Attack Resilience in Object Detectors: Insights from Computing Architecture
Erh-Chung Chen, I-Hsin Chung, Che-Rung Lee |
ACCV (8) | 4 |
| 2024 | Training Sequential CAG Segmentation Models Without Labeled CAG Video DataabstractAlthough sequential Coronary Angiograms (CAG) are often used in cardiac catheterizations, most vessel segmentation techniques are developed just for single images, owing to the expense of collecting and labeling huge amounts of sequential X-Ray images. The CAG segmentation using single images could cause the overlook of critical temporal information among images. Existing video segmentation methods, while they achieve high accuracy for daily life videos, cannot be used directly for sequential CAG, because they are not trained to capture the special structure of coronary articles. In this work, we investigate data manipulation strategies for training sequential CAG segmentation without labeled video data. We leveraged a general-purpose video segmentation model, XMem, for sequential CAG segmentation. We have proposed three data manipulation strategies for model training, including (1) sequence reversal for CAG data; (2) data hybridization for model training; (3) pseudo labels for sequential CAG data. The experimental results show that with good training strategies, one can use general-purpose video segmentation networks for sequential CAG data without explicitly labeled data. The best result of our trained network can achieve an average F1 score of 85.04% and an average region similarity of 73.97%, both of which are higher than the state-of-the-art results. Ting-Chia Yao, Kai-Wen Lin, Chih-Kuo Lee, Po-Hsuan Tseng, Che-Rung Lee |
COMPSAC | 5 |
| 2024 | Overload: Latency Attacks on Object Detection for Edge DevicesabstractNowadays, the deployment of deep learning-based applications is an essential task owing to the increasing demands on intelligent services. In this paper, we investigate latency attacks on deep learning applications. Unlike common adversarial attacks for misclassification, the goal of latency attacks is to increase the inference time, which may stop applications from responding to the requests within a reasonable time. This kind of attack is ubiquitous for various applications, and we use object detection to demonstrate how such kind of attacks work. We also design a framework named Overload to generate latency attacks at scale. Our method is based on a newly formulated optimization problem and a novel technique, called spatial attention. This attack serves to escalate the required computing costs during the inference time, consequently leading to an extended inference time for object detection. It presents a significant threat, especially to systems with limited computing resources. We conducted experiments using YOLOv5 models on Nvidia NX. Compared to existing methods, our method is simpler and more effective. The experimental results show that with latency attacks, the inference time of a single image can be increased ten times longer in reference to the normal setting. Moreover, our findings pose a potential new threat to all object detection tasks requiring non-maximum suppression (NMS), as our attack is NMS-agnostic. Erh-Chung Chen, I-Hsin Chung, Che-Rung Lee |
CVPR | 4 |
| 2024 | LPSD: Low-Rank Plus Sparse Decomposition for Highly Compressed CNN Models
Kuei-Hsiang Huang, Cheng-Yu Sie, Jhong-En Lin, Che-Rung Lee |
PAKDD (1) | 4 |
| 2024 | RPH-PGD: Randomly Projected Hessian for Perturbed Gradient Descent
Chi-Chang Li, Jay Huang, Wing-Kai Hon, Che-Rung Lee |
PAKDD (2) | 4 |
| 2024 | Data filtering for efficient adversarial training
Erh-Chung Chen, Che-Rung Lee |
Pattern Recognit. | 2 |
| 2023 | CoMABO: Covariance Matrix Adaptation for Bayesian OptimizationabstractBayesian optimization, an effective method for searching the optimal solution of black-box functions, usually performs poorly for high-dimensional problems. Although many methods, such as TuRBO (Trust regions Bayesian optimization), have been proposed to solve the over-emphasis of exploration in high-dimensional global acquisition, they may converge slowly or stagnate in some cases. In this paper, we proposed a new method, called CoMABO (Covariance Matrix Adaptation for Bayesian Optimization) to enhance the convergence of TuRBO algorithm. Covariance matrix adaptation is a technique that builds Gaussian models based on the covariance matrix constructed from sampled points. CoMABO utilizes it to explore better candidate points and to optimize the complex surrogate model. Experimental results show CoMABO improves the convergence of TuRBO on various benchmark problems and real world applications. Hsiang-Yu Ku, Che-Rung Lee |
IEEE Big Data | 2 |
| 2023 | GSLAC: GPU Software Level Access Control for Information Isolation on Cloud PlatformsabstractThe massive parallel architecture makes Graphics Processing Unit (GPU) a powerful accelerator for various computational intensive tasks, such as computer games, scientific computation, cryptocurrency, and AI model training and inferences. In many cloud platforms, GPUs are scarce computing resources and shared by multiple users. To achieve information isolation among different user programs, GPU access control is an essential technology to prevent the information leaking for program execution and data access when using GPUs. However, the lack of a zeroing mechanism in GPUs, combined with vulnerabilities in user-land drivers, poses risks to both data confidentiality and system integrity. In this paper, we propose a novel system architecture, called GSLAC, to provide GPU System Level Access Control for information isolation on cloud platforms. GSLAC combines resource isolation and mandatory access control measures with the aim of establishing a secure computing environment. It encompasses an authentication mechanism for authorized GPU access, as well as the integration of mandatory access control mechanisms to safeguard sensitive resources. Furthermore, with a careful design, user programs can be compiled and executed as they do in a normal environment without sacrifying the desired performance. Chia-Chang Li, Po-Cheng Wu, Che-Rung Lee |
CloudCom | 3 |
| 2023 | EFFCA: Enhanced Frangi Filter for Coronary Angiography Segmentation on Mobile Edge DevicesabstractTo better identify the vessels, image segmentation techniques are often applied to coronary angiography (CAG) which reveals the functions and structures of heart's arteries using X-Ray images. Although deep learning based segmentation methods have shown their superiority in accuracy, they are often too complex for medical edge computing, a way to provides prompt diagnoses with minimum hardware cost. In this study, we investigate the method for CAG segmentation on mobile edge devices and propose a novel method, called Enhanced Frangi Filter for Coronary Angiography (EFFCA). Frangi filter is a classical method for vessel segmentation, but suffers from the problems of long processing time for multi-scale search and the vessel breakage problem. EFFCA utilizes a lightweight neural network to recognize the vessel patterns to decide the most suitable scales. It also employs the statistical and connectivity information of vessels to fix the vessel breakage from the segmented results. We have implemented EFFCA on mobile devices to demonstrate its usability. Experimental results show that EFFCA achieves a segmentation accuracy of 95.6% and a specificity of 96.2%, similar to the results of state-of-the-art models. Additionally, EFFCA offers the advantages of a much smaller code size, an efficient training process, and faster inference times on mobile edge devices. Yu-Shuo Wang, Hao-Yun Chen, Chih-Kuo Lee, Che-Rung Lee |
HealthCom | 4 |
| 2023 | Voda: A GPU Scheduling Platform for Elastic Deep Learning in Kubernetes ClustersabstractNowadays, machine learning has become an indispensable service for cloud providers. Elastic training, a novel training paradigm that dynamically adjusts resource allocation for a group of training jobs, has gained popularity due to its ability to effectively utilize accelerators, which are essential for training a massive number of deep learning models. Despite the existence of numerous scheduling algorithms for elastic training, most lack an easy-to-use yet efficient platform to execute them. In this paper, we present Voda, a GPU scheduling platform for elastic deep learning. In contrast to prior approaches that employ parameter servers for elastic training, Voda is designed for AllReduce-style communication, which proves to be more effective, albeit more complex to adjust. Voda, built on top of Kubernetes, consists of a set of loosely coupled components that collect runtime information, dynamically alter the resource allocation, and optimize job placement based on communication costs among underlying GPUs. We implement and compare four scheduling algorithms for elastic training, including three existing methods and one newly proposed, on Voda, with different workloads, job distributions, and arrival patterns. Experimental results demonstrate that no single algorithm dominates all performance metrics, such as average job completion time, running time, or makespan. However, certain algorithms outperform others under specific workloads and job distributions. Additionally, our experiments highlight the significance of job placement in GPU clusters, and our proposed method effectively optimizes communication costs among different workers of a job. Tsung-Tso Hsieh, Che-Rung Lee |
IC2E | 2 |
| 2023 | MontageNet: Annotated Dataset of Furniture Components in Real-World ImagesabstractIndoor understanding is currently a topic that is widely studied in the field of machine learning. Furniture is the most common object in indoor scenes, just as various vehicles are most commonly seen in street scenes. Any object is made up of a combination of functional components. Functional component dismantling and reassembly is an important development for industrial manufacturing to improve efficiency and reduce costs. In the context of understanding indoor scenes, we focus on building a dataset that uses real-world furniture images as part labels. Building a large-scale furniture dataset is very challenging, first of all, the existing dataset has too few real images of furniture, mostly 3D model images, but the diversity of furniture in the real world far exceeds that of 3D models, and real images help improve the calculation speed of model training. Most of the published furniture data is poorly aligned with the annotation data, and even fewer materials perform component segmentation labeling using real furniture images. MontageNet has become a rich resource for part-level 3D shape analysis, semantic understanding, instance segmentation and 3D reconstruction, and other research. Our accompanying empirical studies provide an in-depth analysis of dataset characteristics and performance evaluation of several state-of-the-art methods against our benchmarks. Iuan Kai Fang, Bo-Hao Zhang, Te Lun Liu, Wei Syun Chen, Che-Rung Lee |
MMAsia | 6 |
| 2022 | CSM-DBEN: Container Storage Manager for Data Backup on Edge NodesabstractEdge computing that distributes computing resources close to data sources has become one of the most efficient architectures to solve the latency and the massive connection problems. On top of edge nodes, applications are usually packed in containers to enable fast deployment and multi-tenancy. However, the limited computing resource of edge nodes and application specific containerization make systematic data resilience difficult on edge nodes. In this paper, we present a container storage manager for data backup on edge nodes (CSMDBEN), which provides the live storage backup and recovery for container volumes in the system level of edge nodes. CSM-DBEN leverages virtual machines (VM) to support better security, data isolation, and storage layering. For remote backups, CSM-DBEN performs data encryption and compression for data transmission to enhance the security and bandwidth efficiency. It also dynamically adjusts the data backup frequency, based on the prediction of disk failures, to reduce the system overhead. Experiments show that the incremental data backup time of CSM-DBEN is reduced by 75% $\sim 80$% compared to the commonly used rdiff-backup system. Wei-Cheng Hung, Che-Rung Lee |
CloudCom | 2 |
| 2022 | FLOMD: Fast and Low Overhead Memory Deduplication for Edge NodesabstractEdge computing has become an indispensable architecture to provide real-time services for massive number of IoT devices and 5G platforms. To ease the effort of deploying and managing various applications on edge nodes, people usually pack their applications in containers for edge environment. However, the smooth execution of large amount of containers simultaneously on the nodes with limited computing resources remains a challenge. In this paper, we investigated the memory deduplication technology to relieve the memory pressure on edge nodes. The major difficulty is to quickly discover and merge duplicated memory pages with less CPU consumption. Based on the observed properties of memory usages for containers, we proposed a memory deduplication algorithm, called FLOMD (Fast and Low Overhead Memory Deduplication), which is based on KSM (Kernel Same-page Merging) with four novel techniques: zero page collection, search tree optimization, volatile page identification, and scanning velocity adjustment. Experiments are conducted to compare the sharing efficiency of KSM, UKSM, and FLOMD on edge nodes with various workloads. The results show that the sharing efficiency of FLOMD is more than two times higher than that of others in average. Guann-Ling Shen, Che-Rung Lee |
CloudCom | 2 |
| 2022 | FVMM: Fast VM Migration for Virtualization-based Fault Tolerance Using TemplatesabstractIn the era of cloud computing, virtualization based fault tolerance that utilizes the continuous virtual machine (VM) migration to synchronize a VM and its remote replica is a common technique to achieve high availability. However, traditional live VM migration, whose goal is to minimize the system downtime, has a long duration owing to the expense of the pre-copy for machine status and memory content, which increases the period of failover when failures occur. In this paper, we proposed a new VM migration method, called Fast VM Migration (FVMM), which utilizes the templating technique to accelerate the VM migration by reducing the cost of pre-copy. The templating technique that creates VMs from a master copy, called a template, is a usually used to deploy many similar VMs in a large virtual environment. FVMM employs VM templating to mitigate the cost of pre-copy. Its implementation is optimized with six new techniques: SFVMM and LFVMM, COW templating, continuous templating, asynchronous transmission buffering, VM recovery from continuous templating, and templating regularly. We also applied FVMM for virtualization based fault tolerance to accelerate the failover process. Experimental results show that FVMM is more than 56 times faster than pre-copy live migration for a VM with 16GB, and can reduce upto 32% of the system downtime and the fault tolerance resynchronization time comparing to the methods using pre-copy live migration. Wen-Hsiu Tsai, Po-Jui Tsao, Che-Rung Lee |
CloudCom | 3 |
| 2021 | ADD: A Fine-grained Dynamic Inference Architecture for Semantic Image SegmentationabstractDynamic inference that adaptively skips parts of model execution based on the complexity of input data can effectively reduce the computation cost of deep learning models during the inference. However, current architectures for dynamic inference only consider the exits at the block level, whose results may not be suitable for different applications. In this paper, we present the Auto-Dynamic-DeepLab (ADD), a network architecture that enables the fine-grained dynamic inference for semantic image segmentation. To allow the exit points in the cell level, ADD utilizes Neural Architecture Search (NAS), supported by the framework of Auto-DeepLab, to seek the optimal network structure. In addition, ADD replaces the cells in Auto-DeepLab with the densely connected cells to ease the interference among multiple classifiers and employs the earlier decision maker (EDM) to further improve the performance. Experimental results show that ADD can achieve similar accuracy as Auto-DeepLab in terms of mIoU with a 1.6 times speedup. For the fast mode, ADD can achieve 2.15 times speedup with only a 1.3% accuracy drop compared to those of Auto-DeepLab. Chi-Hsi Kung, Che-Rung Lee |
IROS | 2 |
| 2020 | Towards Fast and Robust Adversarial Training for Image Classification
Erh-Chung Chen, Che-Rung Lee |
ACCV (3) | 2 |
| 2020 | Efficient Group Fault Tolerance for Multi-tier Services in Cloud EnvironmentsabstractFault tolerance is the key technology to achieve high availability for non-stop and long-lasting services, which is usually carried out by the virtualization technology in the era of cloud computing. However, most virtualization-based fault tolerance methods only focus on the resilience of a single server (Individual FT), which can cause great performance degradation for the services that have heavy communication among multiple nodes. One of the solutions is the Group Fault Tolerance (Group FT) technique, which synchronizes a group of VMs within single fault tolerance states to avoid the latency accumulation problem. In this paper, we present an efficient implementation of Group FT, as well as the methods to enhance the performance of Group FT's failover process. Experiments show that Group FT can reduce 88% system latency of the Individual FT when running the OLTP workload in SysBench. Similar results are also shown for a more complicated synthetic multi-tier architecture. Moreover, the performance of the failover process for Group FT is also optimized so that it is comparable to the performance of Individual FT. Chieh-Yu Yu, Che-Rung Lee, Po-Jui Tsao, Yu-Shiang Lin, Tzi-cker Chiueh |
ICC | 2 |
| 2019 | Play as You Like: Timbre-Enhanced Multi-Modal Music Style TransferabstractStyle transfer of polyphonic music recordings is a challenging task when considering the modeling of diverse, imaginative, and reasonable music pieces in the style different from their original one. To achieve this, learning stable multi-modal representations for both domain-variant (i.e., style) and domaininvariant (i.e., content) information of music in an unsupervised manner is critical. In this paper, we propose an unsupervised music style transfer method without the need for parallel data. Besides, to characterize the multi-modal distribution of music pieces, we employ the Multi-modal Unsupervised Image-to-Image Translation (MUNIT) framework in the proposed system. This allows one to generate diverse outputs from the learned latent distributions representing contents and styles. Moreover, to better capture the granularity of sound, such as the perceptual dimensions of timbre and the nuance in instrument-specific performance, cognitively plausible features including mel-frequency cepstral coefficients (MFCC), spectral difference, and spectral envelope, are combined with the widely-used mel-spectrogram into a timbreenhanced multi-channel input representation. The Relativistic average Generative Adversarial Networks (RaGAN) is also utilized to achieve fast convergence and high stability. We conduct experiments on bilateral style transfer tasks among three different genres, namely piano solo, guitar solo, and string quartet. Results demonstrate the advantages of the proposed method in music style transfer with improved sound quality and in allowing users to manipulate the output. Chien-Yu Lu, Min-Xin Xue, Chia-Che Chang, Che-Rung Lee, Li Su 0004 |
AAAI | 4 |
| 2019 | Performance Optimization of SpMV on SparkabstractSparse matrix-vector multiplication (SpMV) is one of the most important computational kernels in solving large scale numerical problems for scientific computing, data analysis, machine learning, and many others. However, its performance optimization on various platforms remains a research problem owing to the diversity of matrix structures and architectural properties. In this paper, we present two performance optimization methods for SpMV on Spark. First, we proposed a new data format, called Block COO plus (BCOO+), which can significantly reduce the number of shuffles in Spark. Second, we designed a new deep convolutional neural network (CNN) to analyze the matrix structure, and automatically choose the right data format for SpMV on Spark. The experimental results show that our method can achieve 3.2 times performance improvement comparing to traditional CSC format. Che-Rung Lee, Feng-Yuan Liu |
IEEE BigData | 2 |
| 2019 | qCUDA: GPGPU Virtualization for High Bandwidth EfficiencyabstractThe increasing demand for machine learning computation contributes to the convergence of high-performance computing and cloud computing, in which the virtualization of Graphics Processing Units (GPUs) becomes a critical issue. Although many GPGPU virtualization frameworks have been proposed, their performance is limited by the bandwidth of data transactions between the virtual machine (VM) and host. In this paper, we present a virtualization framework, qCUDA, to improve the performance of compute unified device architecture (CUDA) programs. qCUDA is based on the virtio framework, providing the para-virtualized driver and the device module for performing the interaction with the API remoting and memory management methods. In our test environment, qCUDA can achieve above 95% of the bandwidth efficiency for most results by comparing it with the native. Also, qCUDA has the features of flexibility and interposition. It can execute CUDA-compatible programs in the Linux and Windows VMs, respectively, on QEMU-KVM hypervisor for GPGPU virtualization. Yu-Shiang Lin, Chun-Yuan Lin, Che-Rung Lee, Yeh-Ching Chung |
CloudCom | 3 |
| 2019 | Performance Optimization for InfiniBand Virtualization on QEMU/KVMabstractThe emergence of machine learning applications has brought the new demands for high performance computing in cloud environment. Besides accelerators, such as GPU or TPU, fast interconnection among and within computers becomes more and more important to achieve efficient training and learning. One of the high-bandwidth, low-latency interconnection architectures is InfiniBand. However, software based virtualization of InfiniBand in QEMU/KVM suffers large performance degradation owing to virtualization overhead and memory allocation problem. In this paper, two techniques, doorbell mapping and memlink, are proposed to optimize the performance the InfiniBand virtualization on QEMU/KVM. Doorbell mapping allows the applications in guest user-space to access the doorbell memory page directly so that the virtualization overhead is minimized. Memlink ensures memory contiguity after virtualization, which is a critical requirement for zero-copy between guest and host. Experiments show that the virtualized InfiniBand with mmap and memlink can achieve near native performance for large data transmissions. Comparing to the previous InfiniBand virtualization on QEMU/KVM, our implementation obtains over 3.5 times performance improvement in various benchmarks. Ming-ting Wei, Yu-Shiang Lin, Che-Rung Lee |
CloudCom | 3 |
| 2019 | Experiences with implementing parallel discrete-event simulation on GPU
Janche Sang, Che-Rung Lee, Vernon Rego, Chung-Ta King |
J. Supercomput. | 2 |
| 2018 | Knowledge Distillation with Feature Maps for Image Classification
Wei-Chun Chen, Chia-Che Chang, Che-Rung Lee |
ACCV (3) | 3 |
| 2018 | Energy Efficient Scheduling for Heterogeneous Fog Computing ArchitecturesabstractHeterogeneous fog computing architectures that hybridize different types of edge nodes can achieve better scalability and lower cost to serve massive number of Internet Of Things (IOT) devices than the centralized architecture of cloud computing. In this paper, we propose an efficient scheduling algorithm to minimize the energy consumption for IOT workflows on heterogeneous fog computing architectures. We first build an integer linear programming model that minimizes the total energy. The purpose of the ILP model is not to be used directly for the computation, but to reveal key factors to minimize the energy consumption in a distributed system. Based on the observations from the model, we derive an energy minimization scheduling (EMS) algorithm that combines different policies to achieve the near optimal scheduling. We implemented and tested our model and algorithm via simulations. The experimental results show that EMS can achieve near optimal energy consumption with much faster execution time. Hsiang-Yi Wu, Che-Rung Lee |
COMPSAC (1) | 2 |
| 2018 | Escaping from Collapsing Modes in a Constrained Space
Chia-Che Chang, Chieh Hubert Lin, Che-Rung Lee, Da-Cheng Juan, Wei Wei 0019, Hwann-Tzong Chen |
ECCV (7) | 3 |
| 2015 | Improving Performance of Convolutional Neural Networks by Separable Filters on GPU
Hao-Ping Kang, Che-Rung Lee |
Euro-Par | 2 |
| 2014 | Taiwan UniCloud: A Cloud Testbed with Collaborative Cloud ServicesabstractThis paper introduces a prototype of Taiwan UniCloud, a community-driven hybrid cloud platform for academics in Taiwan. The goal is to leverage resources in multiple clouds among different organizations. Each self-managing cloud can join the UniCloud platform to share its resources and simultaneously benefit from other clouds with scale-out capabilities. Accordingly, resources are elastic and sharable with each other such as to afford unexpected resource demands to each cloud. The proposed platform provides a web portal to operate each cloud via a uniform user interface. The construction of virtual clusters with multi-core VMs is supplied for parallel and distributed processing models. An object-based storage system is also delivered to federate different storage providers. This paper not only presents the architectural design of Taiwan UniCloud, but also evaluates the performance to demonstrate the possibility of current implementation. Experimental results show the feasibility of the proposed platform as well as the benefit from the cloud federation. Wu-Chun Chung, Po-Chi Shih, Kuan-Chou Lai, Kuanching Li, Che-Rung Lee, Jerry Chou 0001, Ching-Hsien Hsu, Yeh-Ching Chung |
IC2E | 5 |
| 2014 | Achieving cost effective cloud video services via fine grained multicore schedulingabstractCloud computing that possesses highly accessible and elastic computing resources perfectly matches the demands of video services, which employ massive storage and intensive computational power to store, transmit, compress, enhance, and analyze the videos, uploaded from commodity devices and surveillance cameras. However, most existing video processing programs are neither designed to run on parallel environments nor able to efficiently utilize the computational power of cloud platforms, which not only wastes the computing resources but also increases the cost of using cloud platforms. In this paper, we present three strategies to enhance the multicore utilization for video processing, namely producer-consumer model, intra-process overlapping, and inter-process overlapping. We experimented our strategies on a video enhancement program, which performs decoding, dehazing, and encoding, and the results showed the CPU utilization can be improved up to 31% for an 8 core instance, which can significantly reduce the cost in a long run. Hao-Che Kao, Hao-Ping Kang, Che-Rung Lee, Kun-Hsien Lu, Shu-Hsin Chang |
ICPADS | 3 |
| 2014 | Improving GPU Memory Performancewith Artificial Barrier SynchronizationabstractBarrier synchronization, an essential mechanism for a block of threads to guard data consistency, is regarded as a threat to performance. This study, however, provides a different viewpoint for barrier synchronization on GPUs: adding barrier synchronization, even when functionally unnecessary, can improve the performance of some memory-intensive applications. We explain this phenomenon using a memory contention model in which artificial barrier synchronization helps reduce memory contention and preserve data access locality. To yield practical applications, we identify a program pattern: artificial barrier synchronization can be used to synchronize the memory accesses when the data locality among threads is violated. Empirical results from three real-world applications demonstrate that artificial barrier synchronization can increase performance by 10 to 20 percent. Shih-Hsiang Lo, Che-Rung Lee, Quey-Liang Kao, I-Hsin Chung, Yeh-Ching Chung |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2013 | A critical-section-level timing synchronization approach for deterministic multi-core instruction set simulationsabstractThis paper proposes a Critical-Section-Level timing synchronization approach for deterministic Multi-Core Instruction-Set Simulation (MCISS). By synchronizing at each lock access instead of every shared-variable access and using a simple lock usage status managing scheme, our approach significantly improves simulation performance while executing all critical sections in a deterministic order. Experiments show that our approach performs 295% faster than the shared-variable synchronization approach on average and can effectively facilitate system-level software/hardware co-simulation. Fan-Wei Yu, Bo-Han Zeng, Yu-Hung Huang, Hsin-I Wu, Che-Rung Lee, Ren-Song Tsay |
DATE | 5 |
| 2013 | A Compression Algorithm for Fluctuant Data in Smart Grid Database SystemsabstractIn this paper, we present a lossless compression algorithm for fluctuant data, which can be integrated into database system and allows regular database insertion and queries. The algorithm is based on the observation that fluctuant data, although varied violently during small time intervals, have similar patterns over time. The algorithm first partitioned consecutive k records into segments. Those segments are normalized and treated as vectors in k-dimensional space. Classification algorithms are then applied to find representative vectors for those normalized vectors. The classification criterion is that any segments after normalization can find at least one representative vector such that their distance is less than a given threshold. Those representative vectors, called codes, are stored in a codebook. The codebook can be generated offline from a small training dataset, and used repeatedly. The online compression algorithm searches the nearest code for an input segment, and stores only the ID of the code and their difference. Since the difference is small, it can be compressed by Rice coding or Golomb coding.lossless compression algorithm. Chi-Cheng Chuang, Yu-Sheng Chiu, Zhi-Hung Chen, Hao-Ping Kang, Che-Rung Lee |
DCC | 5 |
| 2013 | TLA: Temporal look-ahead processor allocation method for heterogeneous multi-cluster systems
Po-Chi Shih, Kuo-Chan Huang, Che-Rung Lee, I-Hsin Chung, Yeh-Ching Chung |
J. Parallel Distributed Comput. | 3 |
| 2012 | GPU Performance Enhancement via Communication Cost Reduction: Case Studies of Radix Sort and WSN Relay Node Placement ProblemabstractAs the computational power of Graphics Processing Unit (GPU) increases, data transmission becomes the major performance bottleneck. In this study, we investigate two techniques, data streaming and data compression, to reduce the communication cost on GPU. Data streaming enables overlap of communication and computation, whereas data compression reduces the data size transferred among different memory spaces. Although both techniques increase computation cost, overall performance can still be enhanced by reducing communication cost. We demonstrate the effectiveness of the two techniques via two case studies: radix sort and 3-star, a deployment algorithm in wireless sensor networks. For radix sort, a new algorithm, which mixes MSD and LSD algorithms and employs data streaming, is presented. Its performance is 25% faster than the fastest GPU radix sort implementation currently available in the public domain. For the 3-star algorithm, the speed increases several hundreds of times faster than that obtained by the CPU code. The data streaming and data compression, which is a hybrid CPU-GPU algorithm, provide an additional 54% performance improvement to the GPU implementation. Data compression not only reduces communication cost, but also improves the computation time, by which further performance enhancement can be achieved. Che-Rung Lee, Shih-Hsiang Lo, Nan-Hsi Chen, Yeh-Ching Chung, I-Hsin Chung |
CCGRID | 1 |
| 2012 | Accelerating block checkerboard method on GPU for performance enhancement of 2D and 3D Quantum Monte Carlo simulationsabstractQuantum Monte Carlo (QMC) simulations for the recent studies on complex materials were confronted by new computational challenges. Traditional approach to accelerate the simulations by parallel Monte Carlo chains faces serious scalability problems since the speedup is reaching the limitation predicted by Amdahl's law. Fine-grained parallelization of matrix kernels is essential to achieve better performance. In this paper, we investigate the performance optimization techniques on GPU for the most time consuming computational kernel in the Determinant Quantum Monte Carlo (DQMC) simulation: multiplication of matrix exponentials. The matrix, derived from the kinetic Hamiltonian, is highly sparse, and its exponential is approximated by the block checkerboard method, which can represent a matrix exponential as a product of a sequence of sparse matrices. The matrix exponentials from 2D and 3D toruses are focused, and various optimization techniques, such as data streaming and concurrent kernels, are proposed. Experiments show that the proposed optimization techniques can improve the SpMM (Sparse Matrix Multiplication) function, modified from the CUDA SKD SpMV function, up to 16 times and 117 times for 2D and 3D problems respectively. Chi-Cheng Chuang, Yu-Sheng Chiu, Quey-Liang Kao, Zhi-Hung Chen, Che-Rung Lee |
CloudCom | 5 |
| 2012 | Design fast matrix algorithms on high-performance cloud platformsabstractThe low cost and high availability properties of cloud computing make it an attractive solution to not only service-oriented computing, but also to high performance applications. Many of those services and applications entail matrix computations, whose design and performance optimization for cloud platforms haven't been fully studied. Although there is no concrete definition of what a cloud platform should be, especially for high performance purpose, it is believed a heterogeneous computing environment that equips with multicore processors and multiple accelerators is the trend of future high performance cloud architecture. In this paper, a preliminary research on fast matrix multiplication algorithm on multi-GPU is presented, accompanied with a list of future research plans on design fast matrix algorithms on high performance cloud platforms. Quey-Liang Kao, Che-Rung Lee |
CloudCom | 2 |
| 2012 | InfiniBand virtualization on KVMabstractWith the ability to provide on-demand service and to reduce the IT cost, cloud computing becomes more and more popular recently. Virtualization is one of the important technologies in cloud computing, whose main idea is to provide abstractions of the physical resources. However, such abstraction can cause performance degradation, especially for I/O virtualization, which is usually the performance bottleneck in cloud computing. InfiniBand is a network system that provides very low latency (less than 5us) and very high bandwidth (multiple Gbps). Due to its excellent performance, InfiniBand is commonly used in high performance computing (HPC) area. In this paper, we propose Virt-IB for InfiniBand virtualization on Kernel-based Virtual Machine (KVM). The main components of Virt-IB are VM IB library and Virt-IB driver. Our design processes InfiniBand APIs directly in guest VM and communicates with InfiniBand device indirectly to perform the real operations. VM IB library provides API interface and user-level InfiniBand driver. Virt-IB driver provides a channel for VM IB library to write commands into InfiniBand device indirectly. Evaluation results show that our current work is better than network virtualization and it can achieve about 50% performance of native InfiniBand. Yi-Man Ma, Che-Rung Lee, Yeh-Ching Chung |
CloudCom | 2 |
| 2011 | A Parallel Rectangle Intersection Algorithm on GPU+CPUabstractIn this paper, we investigate efficient algorithms and implementations using GPU plus CPU to solve the rectangle intersection problem on a plane. The problem is to report all intersecting pairs of iso-oriented rectangles, whose parallelization on GPUs poses two major computational challenges: data partition and the massive output. The algorithm we presented is called PRI-GC, Parallel Rectangle Intersection algorithm on GPU+CPU, which consists of two phases: mapping and intersection-checking. In the mapping phase, rectangles are hashed into different subspaces (called cells) to reduce the unnecessary intersection checking for far-apart rectangles. In the intersection-checking phase, pairs of rectangles within the same cell are examined in parallel, and the intersecting pairs of rectangles are reported. Several optimization techniques, including rectangles re-ordering, output data compressing/encoding, and the execution overlapping of GPU and CPU, are applied to enhance the performance. We had evaluated the performance of PRI-GC and the result shows over 30x speedup against two well-implemented sequential algorithms on single CPU. The effectiveness of each optimization technique for this problem was evaluated as well. Several parameters, including different degrees of rectangle coverage, different block sizes, and different cell sizes, were also experimented to explore their influences on the performance of PRI-GC. Shih-Hsiang Lo, Che-Rung Lee, Yeh-Ching Chung, I-Hsin Chung |
CCGRID | 2 |
| 2011 | A Performance Goal Oriented Processor Allocation Technique for Centralized Heterogeneous Multi-cluster EnvironmentsabstractThis paper proposes a processor allocation technique named temporal look-ahead processor allocation (TLPA) that makes allocation decision by evaluating the allocation effects on subsequent jobs in the waiting queue. TLPA has two strengths. First, it takes multiple performance factors into account when making allocation decision. Second, it can be used to optimize different performance metrics. To evaluate the performance of TLPA, we compare TLPA with best-fit and fastest-first algorithms. Simulation results show that TLPA has up to 32.75% performance improvement over conventional processor allocation algorithms in terms of average turnaround time in various system configurations. Po-Chi Shih, Kuo-Chan Huang, Che-Rung Lee, I-Hsin Chung, Yeh-Ching Chung |
CCGRID | 3 |
| 2011 | Scalable Communication-Aware Task Mapping Algorithms for Interconnected Multicore SystemsabstractCommunication-aware task mapping algorithms, which map parallel tasks onto processing nodes according to the communication patterns of applications, are essential to reduce the communication time in modern high performance computing. In this paper, we design algorithms specifically for interconnected multicore systems, whose architectural property, namely small number of cores per node, large number of nodes, and large performance gap between the communication within a multicore and among multicores, had brought new challenges and opportunities to the mapping problem. Let k be the number of cores per multicore and n be the number of tasks. We consider the practical case that k ≪ n for k = 2,4, and 6. The designed algorithms are optimal for the mapping measurement, called Maximum Interconnective Message Size (MIMS), and of time complexity merely O(m log m) for m communication pairs. Thus, they are highly scalable for large applications. We had experimented the algorithms on the IBM Blue Gene/P system for two synthetic benchmarks and two applications. The results show good communication performance improvement. I-Hsin Chung, Che-Rung Lee, Jiazheng Zhou, Yeh-Ching Chung |
HPCC | 2 |
| 2011 | On the preconditioner of conjugate gradient method - A power grid simulation perspectiveabstractPreconditioned Conjugate Gradient (PCG) method has been demonstrated to be effective in solving large-scale linear systems for sparse and symmetric positive definite matrices. One critical problem in PCG is to design a good preconditioner, which can significantly reduce the runtime while keeping memory usage efficient. Universal preconditioners are simple and easy to construct, but their effectiveness is highly problem-dependent. On the other hand, domain-specific preconditioners that explore the underlying physical meaning of the matrices usually work better, but are difficult to design. In this paper, we study the problem in the context of power grid simulation, and develop a novel preconditioner based on the power grid structure through simple circuit simulations. Experimental results show 43% reduction in the number of iterations and 23% speedup over existing universal preconditioners. Chung-Han Chou, Nien-Yu Tsai, Hao Yu 0001, Che-Rung Lee, Yiyu Shi 0001, Shih-Chieh Chang 0001 |
ICCAD | 4 |
| 2011 | Redesign of Higher-Level Matrix Algorithms for Multicore and Distributed Architectures and Applications in Quantum Monte Carlo SimulationabstractA matrix operation is referred to as a hard-to-parallel matrix operation (HPMO) if it has serial bottlenecks that are hardly parallelizable. Otherwise, it is referred to as an easy-to-parallel matrix operation (EPMO). Empirical evidences showed the performance scalability of an HPMO is significantly poorer than an EPMO on multicore and distributed architectures. As the result, the design of higher-level algorithms for applications, for the performance considerations on multicore and distributed architectures, should avoid the use of HPMOs as the computational kernels. In this paper, as a case study, we present an HPMO-avoiding algorithm for the Green's function calculation in quantum Monte Carlo simulation. The original algorithm utilizes the QR-decomposition with column pivoting (QRP) as its computational kernel. QRP is an HPMO. The redesigned algorithm maintains the same simulation stability but employs the standard QR decomposition without pivoting (QR), which is an EPMO. Different implementations of the redesigned algorithm on multicore and distributed architectures are investigated. Although some implementations of the redesigned method use about a factor of three more floating-point operations than the original algorithm, they are about 20% faster on a quad core system and 2.5 times faster on a 1024-CPU massively parallel processing system. The broader impact of the redesign of higher-level matrix algorithms to avoid HPMOs in other computational science applications is also discussed. Che-Rung Lee, Zhaojun Bai |
IPDPS | 1 |
| 2010 | Parallelization of DQMC simulation for strongly correlated electron systemsabstractDeterminant Quantum Monte Carlo (DQMC) simulation has been widely used to reveal macroscopic properties of strong correlated materials. However, parallelization of the DQMC simulation is extremely challenging duo to the serial nature of underlying Markov chain and numerical stability issues. We extend previous work with novelty by presenting a hybrid granularity parallelization (HGP) scheme that combines algorithmic and implementation techniques to speed up the DQMC simulation. From coarse-grained parallel Markov chain and task decompositions to fine-grained parallelization methods for matrix computations and Green's function calculations, the HGP scheme explores the parallelism on different levels and maps the underlying algorithms onto different computational components that are suitable for modern high performance heterogeneous computer systems. Practical techniques, such as communication and computation overlapping, message compression and load balancing are also considered in the proposed HGP scheme. We have implemented the DQMC simulation with the HGP scheme on an IBM Blue Gene/P system. The effectiveness of the new scheme is demonstrated through both theoretical analysis and performance results. Experiments have shown over a factor of 80 speedups on an IBM Blue Gene/P system with 1,014 computational processors. Che-Rung Lee, I-Hsin Chung, Zhaojun Bai |
IPDPS | 1 |
| 2008 | Algorithm 879: EIGENTEST - a test matrix generator for large-scale eigenproblemsabstractEigentest is a package that produces real test matrices with known eigensystems. A test matrix, called an eigenmat, is generated in a factored form, in which the user can specify the eigenvalues and has some control over the condition of the eigenvalues and eigenvectors. An eigenmat A of order n requires only O ( n ) storage for its representation. Auxiliary programs permit the computation of ( A − sI ) b , ( A − sI ) T b , ( A − sI ) −1 b , and ( A − sI ) −T b in O ( n ) operations. A special routine computes specified eigenvectors of an eigenmat and the condition of its eigenvalue. Thus eigenmats are suitable for testing algorithms based on Krylov sequences, as well as others based on matrix-vector products. This article introduces the eigenmat and describes implementations in Fortran 77, Fortran 95, C, and Matlab. Che-Rung Lee, G. W. Stewart |
ACM Trans. Math. Softw. | 1 |
| 2000 | A Fast Algorithm for Scheduling Imprecise Computations with Timing Constraints to Minimize Weighted ErrorabstractScheduling tasks with different weights in the imprecise computation model is rather difficult. Each task in the imprecise computation model is logically decomposed into a mandatory subtask and an optional subtask. The mandatory subtask must be completely executed before a deadline to produce an acceptable result; the optional subtask begins after the mandatory subtask to refine the result. The error in the results of a task is measured by the processing time of the unexecuted portion of the optional subtask. This paper proposes a fast algorithm for scheduling imprecise computation with timing constraints on uniprocessor systems. The proposed algorithm can obtain the optimal schedule for different weighted tasks with time complexity O(n log/sup 2/n). Wei-Kuan Shih, Che-Rung Lee, Ching-Hui Tang |
RTSS | 2 |