VLDB 2026 Research / reviewers in the wild / expert
Shanjiang Tang
dblp:67/7742
· DBLP profile ↗
43ranked-venue papers
19as first author
19since 2021 · last 2025
0000-0001-9533-9899ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 11 first-author · 14 since 2021Software engineering, systems software and programming languages · 5 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficiency Optimization Under Spatiotemporal Sharing Fairness for Deep Learning Workloads in Heterogeneous GPU ClustersabstractModern GPU clusters increasingly comprise diverse heterogeneous GPUs, driven by the continuous release of new GPU models. Achieving a balance between fairness and efficiency when scheduling multi-tenant Deep Learning (DL) training jobs on such clusters is inherently challenging. Existing DL training schedulers largely emphasize fairness through GPU temporal sharing, while the spatial dimension of resource allocation is often underexplored. This oversight can lead to GPU fragmentation and suboptimal system performance. In this paper, we propose STS-Fairness, a spatiotemporal sharing fairness scheduler. STS-Fairness partitions each GPU into multiple isolated slots under a novel spatiotemporal fairness constraint and allocates jobs using a round-based allocation mechanism. We guarantee that STS-Fairness achieves overall performance optimality while satisfying spatiotemporal fairness constraints. The scheduling problem is formulated as an integer nonlinear program (INLP) that is solved to optimality in polynomial time via dynamic programming. We deployed the STS-Fairness framework on both physical and simulated heterogeneous clusters and conducted large-scale experiments. These results demonstrate that STS-Fairness reduces average JCT by$1.2 \times$, shortens makespan by$1.24 \times$, and increases throughput by$\mathbf{1. 2 5} \times$compared to state-of-the-art (SoTA) schedulers. Chunhong Du, Mengyu Shi, Shanjiang Tang, Jianhang Tang, Ce Yu, Jian Xiao 0001, Chao Sun 0008, Bin Yang 0043 |
ICPADS | 3 |
| 2025 | Solving online resource-constrained scheduling for follow-up observation in astronomy: A reinforcement learning approach
Ce Yu, Chao Sun 0008, Jizeng Wei, Junhan Ju, Shanjiang Tang |
Future Gener. Comput. Syst. | 6 |
| 2025 | Task Scheduling in Geo-Distributed Computing: A SurveyabstractGeo-distributed computing, a paradigm that assigns computational tasks to globally distributed nodes, has emerged as a promising approach in cloud computing, edge computing, cloud-edge computing, and supercomputer computing (SC). It enables low-latency services, ensures data locality, and handles large-scale applications. As global computing capacity and task demands increase rapidly, scheduling tasks for efficient execution in geo-distributed computing systems has become an increasingly critical research challenge. It arises from the inherent characteristics of geographic distribution, including heterogeneous network conditions, region-specific resource pricing, and varying computational capabilities across locations. Researchers have developed diverse task scheduling methods tailored to geo-distributed scenarios, aiming to achieve objectives such as performance enhancement, fairness assurance, and fault-tolerance improvement. This survey provides a comprehensive and systematic review of task scheduling techniques across four major distributed computing environments, with an in-depth analysis of these approaches based on their core scheduling objectives. Through our analysis, we identify key research challenges and outline promising directions for advancing task scheduling in geo-distributed computing. Yujian Wu, Shanjiang Tang, Ce Yu, Bin Yang 0043, Chao Sun 0008, Jian Xiao 0001, Hutong Wu, Jinghua Feng |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | Accuracy-Efficiency Optimization for Multi-Stage Small Object Detection in Surveillance Video with Collaborative Frame SamplingabstractIn video analytics, accuracy and efficiency are two important metrics and there tend to be a tradeoff between each other. In this paper, we consider accuracy-efficiency optimization for small object detection in surveillance video, which is important and has been widely used in many scenarios such as license plate detection in the traffic domain. Given that small objects tend to being attached to big objects, multi-stage object detection is supposed to be an effective approach to achieve high accuracy for small objects by detecting big objects first and then small objects within the ROIs (Region of interests) of big objects. However, existing studies considered the accuracy-efficiency optimization for small object detection only within the single-stage scenario by changing the frame resolution or sampling rate configuration of video data, which are not suitable for multi-stage detection given that its accuracy-efficiency result is determined by the results of all stages jointly. In this paper, we propose an Adaptive and Collaborative frame Sampling approach named ACS for accuracy-efficiency optimization in the multi-stage small object detection. To improve the efficiency significantly while guaranteeing a given accuracy threshold, ACS dynamically adjusts the sampling rates of all stages collaboratively and periodically using the Karush-Kuhn-Tucker (KKT) condition based on the Lagrangian multiplier method. Additionally, we introduce a tuning knob to allow users to flexibly balance accuracy and efficiency, while ensuring a given accuracy threshold λ. Extensive experiments demonstrate the effectiveness of our approach in improving detection efficiency while guaranteeing diverse accuracy requirements. Chunhong Du, Shanjiang Tang, Song Meng, Jiekai Gou, Ce Yu, Yusen Li, Hao Fu 0021 |
CLUSTER | 2 |
| 2024 | Editorial for the special issue on heterogenous computing
Shanjiang Tang, Yusen Li |
CCF Trans. High Perform. Comput. | 1 |
| 2024 | Fast and accurate novelty detection for large surveillance video
Shanjiang Tang, Ce Yu, Chao Sun 0008, Yusen Li, Jian Xiao 0001 |
CCF Trans. High Perform. Comput. | 1 |
| 2024 | Joint optimization of application placement and resource allocation for enhanced performance in heterogeneous multi-server systems
Pan Lai, Yiran Tao, Yuanai Xie, Shanjiang Tang, Shengquan Liao |
Comput. Networks | 6 |
| 2024 | Fairness-Efficiency Scheduling for Pay-as-You-Go Shared Caching Systems With Long-Term Fairness GuaranteesabstractPay-as-you-go caching systems are now widely used as storage services in cloud computing. However, users’ data caching requirements not only change over time, but also they are affected by workload characteristics, making it difficult to always ensure high efficient use of cache resources. Cache resource sharing is an effective way to improve the efficiency of cache usage. To incentivize users to share caches, it is essential to ensure long-term fairness among multiple users. However, traditional resource allocation strategies canonly guarantee memory less fairness among users,which is not applicable to long-term cache sharing systems. In this paper, we propose a fair allocation policy named FairCache for Pay-as-you-go cache resources. First, FairCache can satisfy four desirable properties of resource allocation: sharing incentive, pay-as-you-usefairness, strategy proofness, and pare to efficiency. Second, FairCache is an efficiency-fairness resource allocation policy based on the efficiency knob θ. The strategy keeps sensitive to the constantly changing cache demands of multiple users within the system by adjusting the efficiency knob θ, thus ensuring long-term multi-user fairness while maximizing the efficiency of cache usage. In addition, FairCache also has an anti-cheating mechanism to avoid possible free-rider problems when multiple users cache access files. Finally, this paper implements the FairCache policy in Alluxio. The experimental results show that FairCache is a lightweight scheduler and that it can maximize the efficiency usage of cache resources while ensuring the long-term for multiple users in the pay-as-you-go Cache systems fairness. Shanjiang Tang, Zhongyu Zhou, Jiekai Gou, Ce Yu, Yusen Li, Hao Fu 0021, Chao Sun 0008, Jian Xiao 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | A Survey on Spark Ecosystem: Big Data Processing Infrastructure, Machine Learning, and Applications (Extended abstract)abstractWith the explosive increase of big data in industry and academic fields, it is important to apply large-scale data processing systems to analyze Big Data. Arguably, Spark is the state-of-the-art in large-scale data computing systems nowadays, due to its good properties including generality, fault tolerance, high performance of in-memory data processing, and scalability. Spark adopts a flexible Resident Distributed Dataset (RDD) programming model with a set of provided transformation and action operators whose operating functions can be customized by users according to their applications. It is originally positioned as a fast and general data processing system. A large body of research efforts have been made to make it more efficient (faster) and general by considering various circumstances since its introduction. In this survey, we aim to have a thorough review of various kinds of optimization techniques on the generality and performance improvement of Spark. We introduce various data management and processing systems, machine learning algorithms and applications supported by Spark. Additionally, we make a discussion on the open issues and challenges for large-scale in-memory data processing with Spark. Shanjiang Tang, Bingsheng He, Ce Yu, Yusen Li, Kun Li 0027 |
ICDE | 1 |
| 2023 | HEGrid: A high efficient multi-channel radio astronomical data gridding framework in heterogeneous computing environments
Ce Yu, Jian Xiao 0001, Shanjiang Tang, Min Long 0001 |
Future Gener. Comput. Syst. | 4 |
| 2022 | EasyNUSC: An Efficient Heterogeneous Computing Framework for Non-uniform Sampling Two-Dimensional Convolution Applications
Ce Yu, Jian Xiao 0001, Hao Fu 0021, Shanjiang Tang, Bo Kang |
ICA3PP | 6 |
| 2022 | Long-Term Fairness Scheduler for Pay-as-You-Use Cache Sharing Systems
Zhongyu Zhou, Shanjiang Tang, Hao Fu 0021, Wanqing Chang, Ce Yu, Chao Sun 0008, Yusen Li, Jian Xiao 0001 |
ICA3PP | 2 |
| 2022 | A method for efficient radio astronomical data gridding on multi-core vector processor
Ce Yu, Jian Xiao 0001, Shanjiang Tang, Hao Fu 0021, Bo Kang, Chenzhou Cui |
Parallel Comput. | 4 |
| 2022 | Reinforcement Learning-Based Resource Partitioning for Improving Responsiveness in Cloud GamingabstractCloud gaming has been very popular in recent years, but issues relating to maintaining low interaction delay to guarantee satisfactory user experience are still prevalent. We observe that the server-side processing delay in cloud gaming system could be heavily influenced by how the resources are partitioned among processes. However, finding the optimal partitioning policy that minimizes the response delay faces several critical challenges. First, fine-grained resource partitioning is non-trivial due to the limitations of hardwre-based resource isolation techniques. Second, game wokload is highly dynamic and unpredictable, making the design of efficient resource partitioning policy more challenging. In this article, we propose an online resource partitioning framework for reducing response delay in cloud gaming, which has several promising properties. First, we divide the processes into disjoint groups and partition resources among process groups, which greatly simplifies the resource partitioning problem while ensuring high partitioning effectiveness. Second, to tackle dynamic workload changes, we classify game workloads into several clusters and maintain separate process grouping plan for each cluster. Third, we leverage reinforcement learning to adaptively choose the best actions for minimizing response delay in real time. We evaluate the proposed framework in a real cloud gaming environment using several real games. The experimental results show that our approach can reduce the response delay by 22 to 41 percent compared to a system without resource partitioning, and outperforms other resource partitioning policies significantly. Yusen Li, Lingjun Pu, Shanjiang Tang, Gang Wang 0001, Xiaoguang Liu 0001 |
IEEE Trans. Computers | 5 |
| 2022 | Fairness-Efficiency Scheduling for Cloud Computing With Soft Fairness GuaranteesabstractFairness and efficiency are two important metrics for users in modern data center computing system. Due to the heterogeneous resource demands of CPU, memory, and network I/O for users’ tasks, it cannot achieve the strict 100 percent fairness and the maximum efficiency at the same time. Existing fairness-efficiency schedulers (e.g., Tetris) can balance such a tradeoff elastically by relaxing fairness constraint for improved efficiency using the knob. However, their approaches areunawareof fairness degradation under different knob configurations, which makes several drawbacks. First, it cannot tell how muchrelaxedfairness can be guaranteed given a knob value. Second, it fails to meet several essential properties such as sharing incentive. To address these issues, we propose a new fairness-efficiency scheduler,QKnober, to balance the fairness and efficiency elastically and flexibly using a tunable fairness knob. QKnober is afairness-sensitivescheduler that can maximize the system efficiency while guaranteeing the$\theta$-soft fairness by modeling the whole allocation as a combination offairness-orientedallocation andefficiency-orientedallocation. Moreover, QKnober satisfies fairness properties of sharing incentive, envy-freeness and pareto efficiency given a proper knob value. We have implemented QKnober in YARN and evaluated it using both testbed and simulated experiments. The results show that QKnober outperforms its alternatives DRF and Tetris by 31.2 and 4.5 percent, respectively. Shanjiang Tang, Ce Yu, Yusen Li |
IEEE Trans. Cloud Comput. | 1 |
| 2022 | A Survey on Spark Ecosystem: Big Data Processing Infrastructure, Machine Learning, and ApplicationsabstractWith the explosive increase of big data in industry and academic fields, it is important to apply large-scale data processing systems to analyze Big Data. Arguably, Spark is the state-of-the-art in large-scale data computing systems nowadays, due to its good properties including generality, fault tolerance, high performance of in-memory data processing, and scalability. Spark adopts a flexible Resident Distributed Dataset (RDD) programming model with a set of provided transformation and action operators whose operating functions can be customized by users according to their applications. It is originally positioned as afastandgeneraldata processing system. A large body of research efforts have been made to make it more efficient (faster) and general by considering various circumstances since its introduction. In this survey, we aim to have a thorough review of various kinds of optimization techniques on the generality and performance improvement of Spark. We introduce Spark programming model and computing system, discuss the pros and cons of Spark, and have an investigation and classification of various solving techniques in the literature. Moreover, we also introduce various data management and processing systems, machine learning algorithms and applications supported by Spark. Finally, we make a discussion on the open issues and challenges for large-scale in-memory data processing with Spark. Shanjiang Tang, Bingsheng He, Ce Yu, Yusen Li, Kun Li 0027 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | DVQShare: An Analytics System for DNN-based Video QueriesabstractApplying deep neural networks (DNNs) to video analytics tasks has drawn attention from both academic and industry communities. However, due to the high computational complexity of DNN models and the explosion of video data, it is challenging to process massive concurrent video queries efficiently and effectively. In this paper, we propose a video analytics system named DVQShare to process DNN-based video queries in a batch mode. The key idea is sharing, including time sharing, spatial sharing, and logical sharing. In principle, sharing across queries can help us reduce the overall amount of frames to be analyzed, which can help us improve the overall performance and reduce the monetary cost. Two modules are designed to process video queries by exploiting the above three sharing opportunities. First, an analysis module is integrated to guide the generation of query processing plans. Within this module, temporal sharing is considered to reuse historical results produced by other queries to remove pending frames that have been analyzed, and spatial sharing is adopted to avoid redundant processing over overlapping video clips. Additionally, we utilize logical sharing to further improve system's overall performance by considering the logical relationship between queries. Second, a query processing engine is devised to execute the query pipeline generated by the analysis module and return the final results. In experiments, we implement a prototype of the DVQShare system based on MXNet, and results show that it can achieve up to 2X performance speedup. Hao Fu 0021, Shanjiang Tang, Ce Yu, Yusen Li, Yanjie Liu |
CCGRID | 2 |
| 2021 | EDL-COVID: Ensemble Deep Learning for COVID-19 Case Detection From Chest X-Ray ImagesabstractEffective screening of COVID-19 cases has been becoming extremely important to mitigate and stop the quick spread of the disease during the current period of COVID-19 pandemic worldwide. In this article, we consider radiology examination of using chest X-ray images, which is among the effective screening approaches for COVID-19 case detection. Given deep learning is an effective tool and framework for image analysis, there have been lots of studies for COVID-19 case detection by training deep learning models with X-ray images. Although some of them report good prediction results, their proposed deep learning models might suffer from overfitting, high variance, and generalization errors caused by noise and a limited number of datasets. Considering ensemble learning can overcome the shortcomings of deep learning by making predictions with multiple models instead of a single model, we proposeEDL-COVID, an ensemble deep learning model employing deep learning and ensemble learning. The EDL-COVID model is generated by combining multiple snapshot models of COVID-Net, which has pioneered in an open-sourced COVID-19 case detection method with deep neural network processed chest X-ray images, by employing a proposed weighted averaging ensembling method that is aware of different sensitivities of deep learning models on different classes types. Experimental results show that EDL-COVID offers promising results for COVID-19 case detection with an accuracy of 95%, better than COVID-Net of 93.3%. Shanjiang Tang, Chunjiang Wang, Jiangtian Nie, Neeraj Kumar 0001, Yang Zhang 0025, Zehui Xiong, Ahmed Barnawi |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | HGP4CNN: an efficient parallelization framework for training convolutional neural networks on modern GPUs
Hao Fu 0021, Shanjiang Tang, Bingsheng He, Ce Yu |
J. Supercomput. | 2 |
| 2020 | Deep Learning Assisted Resource Partitioning for Improving Performance on Commodity ServersabstractIn this paper, we introduce a deep reinforcement learning (DRL) framework for solving the problem of partitioning LLC and memory bandwidth coordinately in an end-to-end manner. To this end, we formulate the problem as a markov decision process and utilize DRL algorithm to derive the optimal partition. To avoid the extensive cost of training the policy on physical server, we present a model-based solution, where a reward prediction model is leveraged to train the partitioning policy offline. To construct a precise reward prediction model, we introduce a novel representation for the partitioning scheme, where graph convolutional networks (GCN) is employed to represent the LLC partition as a bipartite graph so that those heterogeneous but identical partitions could result in the same representations and thus eases the prediction task. Ruobing Chen 0002, Jinping Wu, Haosen Shi 0001, Yusen Li, Haiyan Yin, Shanjiang Tang, Xiaoguang Liu 0001, Gang Wang 0001 |
PACT | 6 |
| 2020 | Balancing Fairness and Efficiency for Cache Sharing in Semi-external Memory SystemabstractData caching and sharing is an effective approach for achieving high performance to many applications in shared platforms such as the cloud. DRAM and SSD are two popular caching devices widely used by many large-scale data application systems such Hadoop and Spark. Due to the limited size of DRAM as well as the large access latency of SSD (relative to DRAM), there is a trend of integrating DRAM and SSD (called semi-external memory) together for large-scale data caching. Shanjiang Tang, Qifei Chai, Ce Yu, Yusen Li, Chao Sun 0008 |
ICPP | 1 |
| 2019 | GAugur: Quantifying Performance Interference of Colocated Games for Improving Resource Utilization in Cloud GamingabstractCloud gaming has been very popular recently, but providing satisfactory gaming experiences to players at a modest cost is still challenging. Colocating several games onto one server could improve server utilization. To enable efficient colocations while providing Quality of Service (QoS) guarantees, a precise quantification of performance interference among colocated games is required. However, achieving such precise interference prediction is very challenging for games due to the complexity introduced by the contention on many shared resources across CPU and GPU. Moreover, the distinctive properties of cloud gaming require that the prediction model should be constructed beforehand and the prediction should be made instantaneously at request arrivals, which further increases the difficulty. The existing solutions are either not applicable or not effective due to many limitations. In this paper, we present GAugur, a novel methodology that enables highly accurate prediction of the performance interference among games arbitrarily colocated. By leveraging machine learning technologies, GAugur is able to capture the complex relationship between the interference and the contention features of colocated games. We evaluate GAugur through extensive experiments using a large number of real popular games. The results show that GAugur is able to identify whether a colocated game satisfies QoS requirement within an average error of 5%, and is able to quantify the performance degradation of a colocated game within an average error of 7.9%, which significantly outperforms the alternatives. Moreover, GAugur incurs an offline profiling cost linear to the number of games, and negligible overhead for online prediction. We apply GAugur to guiding efficient game colocations for cloud gaming. Experimental results show that GAugur is able to increase the resource utilization by 20% to 60%, and improve the overall performance by up to 15%, compared to the state-of-the-art solutions. Yusen Li, Chuxu Shan, Ruobing Chen 0002, Xueyan Tang, Wentong Cai 0001, Shanjiang Tang, Xiaoguang Liu 0001, Gang Wang 0001, Xiaoli Gong, Ying Zhang 0015 |
HPDC | 6 |
| 2019 | HDF5-Based I/O Optimization for Extragalactic HI Data Pipeline of FAST
Yiming Ji, Ce Yu, Jian Xiao 0001, Shanjiang Tang |
ICA3PP (2) | 4 |
| 2019 | Themis: Efficient and Adaptive Resource Partitioning for Reducing Response Delay in Cloud GamingabstractCloud gaming has been increasing in popularity recently, but issues relating to maintaining low interaction delay for users to guarantee satisfactory gaming experience is still prevalent. Interaction delays caused by server-side processing are heavily influenced by how the processes partition the resources. However, finding the optimal partitioning policy that minimizes the response delay is complicated by several critical challenges. In this paper, we propose Themis, a system that enables efficient and adaptive online resource partitioning for reducing response delay in cloud gaming. Briefly, Themis employs machine learning technology to build a performance model which is able to capture the complex relationships between resource partition and system performance. With this model, Themis divides the processes into disjoint groups and partitions resources among process groups, which greatly simplifies the resource partition problem while ensuring high partitioning effectiveness. To tackle dynamic workload changes, Themis leverages reinforcement learning to learn how different partitioning actions affect system performance in an online manner, and adaptively choose the best actions for minimizing response delay in real time. We evaluate Themis in a real cloud gaming environment using several real games. The experimental results show that Themis can reduce the response delay by 17% to 36% compared to a system without resource partitioning, and outperforms other resource partitioning policies significantly. To the best of our knowledge, this is the first work to optimize response delay in cloud gaming through resource partitioning. Yusen Li, Lingjun Pu, Trent Marbach, Shanjiang Tang, Gang Wang 0001, Xiaoguang Liu 0001 |
ACM Multimedia | 6 |
| 2019 | An Adaptive Efficiency-Fairness Meta-Scheduler for Data-Intensive ComputingabstractIn data-intensive cluster computing platforms such as Hadoop YARN, efficiency and fairness are two important factors for system design and optimizations. Previous studies are either for efficiency or for fairness solely, without considering the tradeoff between efficiency and fairness. Recent studies observe that there is a tradeoff between efficiency and fairness because of resource contention between users/jobs. By leveraging the existing schedulers, a meta-scheduler is able to dynamically choose one of them for job/task scheduling at runtime. In this paper, we propose a meta-scheduler called FLEX to realize the tradeoff between system efficiency and fairness in Hadoop YARN. FLEX combines multiple existing schedulers into a single aggregated view without any modification on the original schedulers. Equipped with these candidate schedulers, FLEX utilizes machine learning approach to adaptively choose the most proper scheduler according to the characteristic of current running workload and user-defined Service Level Agreement (SLA). We implement FLEX in Hadoop YARN. We conduct experiments with real deployment in a local cluster and perform simulation studies with production traces. Experimental results show that the FLEX outperforms the state-of-the-art approach in two aspects: 1) Given a predefined threshold on the fairness loss, the FLEX reduces the makespan by up to 22 and 24 percent in real deployment and the large-scale simulation, respectively; 2) Given the predefined threshold on the makespan reduction, the FLEX reduces the fairness loss by up to 75 and 73 percent in real deployment and the large-scale simulation, respectively. Zhaojie Niu, Shanjiang Tang, Bingsheng He |
IEEE Trans. Serv. Comput. | 2 |
| 2018 | An Efficient Retrieval Method for Astronomical Catalog Time Series Data
Bingyao Li 0001, Ce Yu, Xiaoteng Hu, Jian Xiao 0001, Shanjiang Tang, Lianmeng Li, Bin Ma 0022 |
ICA3PP (1) | 5 |
| 2018 | Correction to: An Efficient Retrieval Method for Astronomical Catalog Time Series Data
Bingyao Li 0001, Ce Yu, Xiaoteng Hu, Jian Xiao 0001, Shanjiang Tang, Lianmeng Li, Bin Ma 0022 |
ICA3PP (1) | 5 |
| 2018 | GpDL: A Spatially Aggregated Data Layout for Long-Term Astronomical Observation Archive
Ce Yu, Chao Sun 0008, Shanjiang Tang, Xiangfei Meng |
ICA3PP (2) | 4 |
| 2018 | GLP4NN: A Convergence-invariant and Network-agnostic Light-weight Parallelization Framework for Deep Neural Networks on Modern GPUsabstractIn this paper, we propose a network-agnostic and convergence-invariant light-weight parallelization framework, namely GLP4NN, to accelerate the training of Deep Neural Networks (DNNs) by taking advantage of emerging GPU features, especially concurrent kernel execution. To determine the number of concurrent kernels on the fly, we design an analytical model in the kernel analyzer module and integrate a compact asynchronous resource tracker in the resource tracker module for collecting runtime configurations of kernels with low memory and time overheads. We further develop a runtime scheduler module and a pool-based stream manager for handling GPU work queues in GLP4NN to avoid consuming too many CPU threads or processes while dispatching workloads to GPU devices. In our experiments, we integrate GLP4NN into Caffe to accelerate the batch-based training of four well-known networks on NVIDIA GPUs. Experimental results show GLP4NN is able to achieve a speedup of up to 4X over the original implementation as well as keep the convergence property of networks. Hao Fu 0021, Shanjiang Tang, Bingsheng He, Ce Yu |
ICPP | 2 |
| 2018 | QKnober: A Knob-Based Fairness-Efficiency Scheduler for Cloud Computing with QoS Guarantees
Shanjiang Tang, Ce Yu, Chao Sun 0008, Jian Xiao 0001, Yinglong Li |
ICSOC | 1 |
| 2018 | Long-Term Multi-Resource Fairness for Pay-as-you Use Computing SystemsabstractMany current computing systems such as clouds and supercomputers charge users for their resource usages. A user's demand is often changing over time, indicating that it is difficult to keep the high resource utilization all the time for cost efficiency. Resource sharing is a classical and effective approach for high resource utilization. In view of the heterogeneous resource demands of users' workloads, multi-resource allocation fairness is a must for resource sharing in such pay-as-you-use computing systems. However, we find that, existing multi-resource fair policies such as Dominant Resource Fairness (DRF), implemented in currently popular resource management systems such as Apache YARN [4] and Mesos [23], are not suitable for the pay-as-you-use computing systems. We show that this is because of their memoryless characteristic that can cause the following problems in the pay-as-you-use computing systems: 1). users can get resource benefits by cheating; 2). users might not be able to get the total amount of resources that they are entitled to in terms of their resource contributions. In this paper, we propose a new policy called H-MRF, which generalizes DRF and Asset Fairness with the long-term notion. We show that it can address these problems and is suitable for pay-as-you-use computing systems. We have implemented it into YARN by developing a prototype called MRYARN. Finally, we evaluate H-MRF using both testbed and simulated experiments. The experimental results show that there are about 1.1 ~1.5 sharing benefit degrees and 1.2× ~ 1.8× performance improvement for users with H-MRF, better than existing fair schedulers. Shanjiang Tang, Zhaojie Niu, Bingsheng He, Bu-Sung Lee, Ce Yu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | Fair Resource Allocation for Data-Intensive Computing in the CloudabstractTo address the computing challenge of `big data', a number of data-intensive computing frameworks (e.g., MapReduce, Dryad, Storm and Spark) have emerged and become popular. YARN is a de facto resource management platform that enables these frameworks running together in a shared system. However, we observe that, in cloud computing environment, the fair resource allocation policy implemented in YARN is not suitable because of its memoryless resource allocation fashion leading to violations of a number of good properties in shared computing systems. This paper attempts to address these problems for YARN. Both single-level and hierarchical resource allocations are considered. For single-level resource allocation, we propose a novel fair resource allocation mechanism called Long-Term Resource Fairness (LTRF)for such computing. For hierarchical resource allocation, we propose Hierarchical Long-Term Resource Fairness (H-LTRF) by extending LTRF. We show that both LTRF and H-LTRF can address these fairness problems of current resource allocation policy and are thus suitable for cloud computing. Finally, we have developed LTYARN by implementing LTRF and H-LTRF in YARN, and our experiments show that it leads to a better resource fairness than existing fair schedulers of YARN. Shanjiang Tang, Bu-Sung Lee, Bingsheng He |
IEEE Trans. Serv. Comput. | 1 |
| 2017 | CMSA: a heterogeneous CPU/GPU computing system for multiple similar RNA/DNA sequence alignmentabstractThe multiple sequence alignment (MSA) is a classic and powerful technique for sequence analysis in bioinformatics. With the rapid growth of biological datasets, MSA parallelization becomes necessary to keep its running time in an acceptable level. Although there are a lot of work on MSA problems, their approaches are either insufficient or contain some implicit assumptions that limit the generality of usage. First, the information of users’ sequences, including the sizes of datasets and the lengths of sequences, can be of arbitrary values and are generally unknown before submitted, which are unfortunately ignored by previous work. Second, the center star strategy is suited for aligning similar sequences. But its first stage, center sequence selection, is highly time-consuming and requires further optimization. Moreover, given the heterogeneous CPU/GPU platform, prior studies consider the MSA parallelization on GPU devices only, making the CPUs idle during the computation. Co-run computation, however, can maximize the utilization of the computing resources by enabling the workload computation on both CPU and GPU simultaneously. This paper presents CMSA, a robust and efficient MSA system for large-scale datasets on the heterogeneous CPU/GPU platform. It performs and optimizes multiple sequence alignment automatically for users’ submitted sequences without any assumptions. CMSA adopts the co-run computation model so that both CPU and GPU devices are fully utilized. Moreover, CMSA proposes an improved center star strategy that reduces the time complexity of its center sequence selection process from O(mn 2) to O(mn). The experimental results show that CMSA achieves an up to 11× speedup and outperforms the state-of-the-art software. CMSA focuses on the multiple similar RNA/DNA sequence alignment and proposes a novel bitmap based algorithm to improve the center star strategy. We can conclude that harvesting the high performance of modern GPU is a promising approach to accelerate multiple sequence alignment. Besides, adopting the co-run computation model can maximize the entire system utilization significantly. The source code is available at https://github.com/wangvsa/CMSA . Chen Wang 0004, Shanjiang Tang, Ce Yu, Quan Zou 0001 |
BMC Bioinform. | 3 |
| 2016 | Elastic multi-resource fairness: balancing fairness and efficiency in coupled CPU-GPU architecturesabstractFairness and efficiency are two important concerns for users in a shared computer system, and there tends to be a tradeoff between them. Heterogeneous computing poses new challenging issues on the fair allocation of computational resources among users due to the availability of different kinds of computing devices (e.g., CPU and GPU). Prior work either considers the fair resource allocation separately for each computing device or is unable to balance flexibly the tradeoff between the fairness and system utilization. In this work, we consider an emerging heterogeneous computing system with coupled CPU and GPU into a single chip. We first show that it is essential to have a new fair policy for coupled CPU-GPU architectures that is capable of considering both the CPU and the GPU as a whole in fair resource allocation and being aware of the system utilization maximization. We then propose a fair policy called Elastic Multi-Resource Fairness (EMRF) for coupled CPU-GPU architectures, by modeling CPU and GPU as two resource types and viewing the resource fairness problem as a multi-resource fairness problem. It extends DRF by adding a knob that allows users to tune and balance fairness and performance flexibly, and considers the fair allocation of computational resources as a whole for CPU and GPU devices. We show that EMRF satisfies fairness properties of sharing incentive, envy-freeness and pareto efficiency. Finally, we evaluate EMRF using real experiments, and the results show that EMRF can achieve better performance and fairness. Shanjiang Tang, Bingsheng He, Shuhao Zhang 0001, Zhaojie Niu |
SC | 1 |
| 2016 | A general and fast distributed system for large-scale dynamic programming applications
Chen Wang 0004, Ce Yu, Shanjiang Tang, Jian Xiao 0001, Xiangfei Meng |
Parallel Comput. | 3 |
| 2016 | Dynamic Job Ordering and Slot Configurations for MapReduce WorkloadsabstractMapReduce is a popular parallel computing paradigm for large-scale data processing in clusters and data centers. A MapReduce workload generally contains a set of jobs, each of which consists of multiple map tasks followed by multiple reduce tasks. Due to 1) that map tasks can only run in map slots and reduce tasks can only run in reduce slots, and 2) the general execution constraints that map tasks are executed before reduce tasks, different job execution orders and map/reduce slot configurations for a MapReduce workload have significantly different performance and system utilization. This paper proposes two classes of algorithms to minimize the makespan and the total completion time for an offline MapReduce workload. Our first class of algorithms focuses on the job ordering optimization for a MapReduce workload under a given map/reduce slot configuration. In contrast, our second class of algorithms considers the scenario that we can perform optimization for map/reduce slot configuration for a MapReduce workload. We perform simulations as well as experiments on Amazon EC2 and show that our proposed algorithms produce results that are up to 15 ~ 80 percent better than currently unoptimized Hadoop, leading to significant reductions in running time in practice. Shanjiang Tang, Bu-Sung Lee, Bingsheng He |
IEEE Trans. Serv. Comput. | 1 |
| 2015 | Gemini: An Adaptive Performance-Fairness Scheduler for Data-Intensive Cluster ComputingabstractIn data-intensive cluster computing platforms such as Hadoop YARN, performance and fairness are two important factors for system design and optimizations. Many previous studies are either for performance or for fairness solely, without considering the tradeoff between performance and fairness. Recent studies observe that there is a tradeoff between performance and fairness because of resource contention between users/jobs. However, their scheduling algorithms for bi-criteria optimization between performance and fairness are static, without considering the impact of different workload characteristics on the tradeoff between performance and fairness. In this paper, we propose an adaptive scheduler called Gemini for Hadoop YARN. We first develop a model with the regression approach to estimate the performance improvement and the fairness loss under the sharing computation compared to the exclusive non-sharing scenario. Next, we leverage the model to guide the resource allocation for pending tasks to optimize the performance of the cluster given the user-defined fairness level. Instead of using a static scheduling policy, Gemini adaptively decides the proper scheduling policy according to the current running workload. We implement Gemini in Hadoop YARN. Experimental results show that Gemini outperforms the state-of-the-art approach in two aspects. 1) For the same fairness loss, Gemini improves the performance by up to 225% and 200% in real deployment and the large-scale simulation, respectively, 2) For the same performance improvement, Gemini reduces the fairness loss up to 70% and 62.5% in real deployment and the large-scale simulation, respectively. Zhaojie Niu, Shanjiang Tang, Bingsheng He |
CloudCom | 2 |
| 2014 | Towards Economic Fairness for Big Data Processing in Pay-as-You-Go Cloud ComputingabstractRecent trends indicate that the pay-as-you-go Infrastructure-as-a-Service (IaaS) cloud computing has become a popular platform for big data processing applications, due to its merits of accessibility, elasticity and flexibility. However, the resource demands of processing workloads are often varying over time for individual users, implying that it is hard for a user to keep the high resource utilization for cost efficiency all the time. Resource sharing is a classic and effective approach to improve the resource utilization via consolidating multiple users' workloads. However, we show that, current existing fair policies such as max-min fairness, widely adopted and implemented in many popular big data processing systems including YARN, Spark, Mesos, and Dryad, are not suitable for pay-as-you-go cloud computing. We show that it is because of their memory less allocation feature which can arise a series of problems in the pay-as-you-go cloud environment, namely, cost-inefficient workload submission, untruthfulness and resource-as-you-pay unfairness. This paper presents these problems and outlines our plans to address them for pay-as-you-go cloud computing. We introduce our preliminary work done on the single-resource fairness and our ongoing work for multi-resource fairness, and outline our future work. Shanjiang Tang, Bu-Sung Lee, Bingsheng He |
CloudCom | 1 |
| 2014 | Long-term resource fairness: towards economic fairness on pay-as-you-use computing systemsabstractFair resource allocation is a key building block of any shared computing system. However, MemoryLess Resource Fairness (MLRF), widely used in many existing frameworks such as YARN, Mesos and Dryad, is not suitable for pay-as-you-use computing. To address this problem, this paper proposes Long-Term Resource Fairness (LTRF), a novel fair resource allocation mechanism. We show that LTRF satisfies several highly desirable properties. First, LTRF incentivizes clients to share resources via group-buying by ensuring that no client is better off in a computing system that she buys and uses individually. Second, LTRF incentivizes clients to submit non-trivial workloads and be willing to yield unneeded resources to others. Third, LTRF has a resource-as-you-pay fairness property, which ensures the amount of resources that each client should get according to her monetary cost, despite that her resource demand varies over time. Finally, LTRF is strategy-proof, since it can make sure that a client cannot get more resources by lying about her demand. We have implemented LTRF in YARN by developing LTYARN, a long-term YARN fair scheduler, and shown that it leads to a better resource fairness than other state-of-the-art fair schedulers. Shanjiang Tang, Bu-Sung Lee, Bingsheng He, Haikun Liu |
ICS | 1 |
| 2014 | DynamicMR: A Dynamic Slot Allocation Optimization Framework for MapReduce ClustersabstractMapReduce is a popular computing paradigm for large-scale data processing in cloud computing. However, the slot-based MapReduce system (e.g., Hadoop MRv1) can suffer from poor performance due to its unoptimized resource allocation. To address it, this paper identifies and optimizes the resource allocation from three key aspects. First, due to the pre-configuration of distinct map slots and reduce slots which are not fungible, slots can be severely under-utilized. Because map slots might be fully utilized while reduce slots are empty, and vice-versa. We propose an alternative technique called Dynamic Hadoop SlotAllocation by keeping the slot-based model. It relaxes the slot allocation constraint to allow slots to be reallocated to either map or reduce tasks depending on their needs. Second, the speculative execution can tackle the straggler problem, which has shown to improve the performance for a single job but at the expense of the cluster efficiency. In view of this, we propose Speculative Execution Performance Balancing to balance the performance tradeoff between a single job and a batch of jobs. Third, delay scheduling has shown to improve the data locality but at the cost of fairness. Alternatively, we propose a technique called Slot PreSchedulingthat can improve the data locality but with no impact on fairness. Finally, by combining these techniques together, we form a step-by-step slot allocation system called DynamicMR that can improve the performance of MapReduce workloads substantially. The experimental results show that our DynamicMR can improve the performance of Hadoop MRv1 significantly while maintaining the fairness, by up to 46~115 percent for single jobs and 49~112 percent for multiple jobs. Moreover, we make a comparison with YARN experimentally, showing that DynamicMR outperforms YARN by about 2~9 percent for multiple jobs due to its ratio control mechanism of running map/reduce tasks. Shanjiang Tang, Bu-Sung Lee, Bingsheng He |
IEEE Trans. Cloud Comput. | 1 |
| 2013 | Dynamic slot allocation technique for MapReduce clustersabstractMapReduce is a popular parallel computing paradigm for large-scale data processing in clusters and data centers. However, the slot utilization can be low, especially when Hadoop Fair Scheduler is used, due to the pre-allocation of slots among map and reduce tasks, and the order that map tasks followed by reduce tasks in a typical MapReduce environment. To address this problem, we propose to allow slots to be dynamically (re)allocated to either map or reduce tasks depending on their actual requirement. Specifically, we have proposed two types of Dynamic Hadoop Fair Scheduler (DHFS), for two different levels of fairness (i.e., cluster and pool level). The experimental results show that the proposed DHFS can improve the system performance significantly (by 32% ~ 55% for a single job and 44% ~ 68% for multiple jobs) while guaranteeing the fairness. Shanjiang Tang, Bu-Sung Lee, Bingsheng He |
CLUSTER | 1 |
| 2013 | MROrder: Flexible Job Ordering Optimization for Online MapReduce Workloads
Shanjiang Tang, Bu-Sung Lee, Bingsheng He |
Euro-Par | 1 |
| 2012 | EasyPDP: An Efficient Parallel Dynamic Programming Runtime System for Computational BiologyabstractDynamic programming (DP) is a popular and efficient technique in many scientific applications such as computational biology. Nevertheless, its performance is limited due to the burgeoning volume of scientific data, and parallelism is necessary and crucial to keep the computation time at acceptable levels. The intrinsically strong data dependency of dynamic programming makes it difficult and error-prone for the programmer to write a correct and efficient parallel program. Therefore, this paper builds a runtime system named EasyPDP aiming at parallelizing dynamic programming algorithms on multicore and multiprocessor platforms. Under the concept of software reusability and complexity reduction of parallel programming, a DAG Data Driven Model is proposed, which supports those applications with a strong data interdependence relationship. Based on the model, EasyPDP runtime system is designed and implemented. It automatically handles thread creation, dynamic data task allocation and scheduling, data partitioning, and fault tolerance. Five frequently used DAG patterns from biological dynamic programming algorithms have been put into the DAG pattern library of EasyPDP, so that the programmer can choose to use any of them according to his/her specific application. Besides, an ideal computing distribution model is proposed to discuss the optimal values for the performance tuning arguments of EasyPDP. We evaluate the performance potential and fault tolerance feature of EasyPDP in multicore system. We also compare EasyPDP with other methods such as Block-Cycle Wavefront (BCW). The experimental results illustrate that EasyPDP system is fine and provides an efficient infrastructure for dynamic programming algorithms. Shanjiang Tang, Ce Yu, Bu-Sung Lee, Huabei Wu |
IEEE Trans. Parallel Distributed Syst. | 1 |