EDBT 2026 Demo / reviewers in the wild / expert
Ce Yu
dblp:47/4350
· DBLP profile ↗
59ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0003-2416-4547ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 44 · 3 first-author · 14 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Power load forecasting based on time-frequency domain feature fusion
Wenhua Jiao, Ce Yu, Feiyi Xu |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | OFMAD-TC: A tropical cyclone detection method with optical flow and morphology awareness
Xiaoxian Tian, Lu Yang 0007, Chongke Bi, Ce Yu |
Neurocomputing | 4 |
| 2026 | A Hierarchical Conflict Resolution Framework With Graph Transformer-Based Reinforcement Learning for Heterogeneous UAV NetworksabstractDense unmanned aerial vehicle (UAV) networks comprising different types of UAVs have been widely applied across various domains, leading to potential conflicts within heterogeneous UAV networks. To ensure UAV safety, effective conflict resolution is crucial. However, existing learning-based methods are primarily designed for homogeneous UAVs and struggle to learn effective maneuvers for UAVs with varying attributes (like velocity and safety radius) or diverse motion pattern, such as hovering. To address these challenge, we propose a hierarchical framework with graph transformer-based reinforcement learning (HGT-RL) for conflict resolution in heterogeneous UAV networks. Specifically, a novel heterogeneous graph transformer network is employed to enhance the extraction of heterogeneous node information through orderly embeddings while incorporating a graph transformer network. To reduce the training complexity for heterogeneous UAVs, we generate the heuristic maneuvers at each time step, with the reinforcement learning model trained solely to refine these initial actions or perform emergency hovering. Experimental results demonstrate that HGT-RL outperforms previous methods in heterogeneous UAV networks, and we further validate its scalability across multiple scenarios. Ce Yu, Wenbo Du 0001 |
IEEE Internet Things J. | 3 |
| 2025 | Efficiency Optimization Under Spatiotemporal Sharing Fairness for Deep Learning Workloads in Heterogeneous GPU ClustersabstractModern GPU clusters increasingly comprise diverse heterogeneous GPUs, driven by the continuous release of new GPU models. Achieving a balance between fairness and efficiency when scheduling multi-tenant Deep Learning (DL) training jobs on such clusters is inherently challenging. Existing DL training schedulers largely emphasize fairness through GPU temporal sharing, while the spatial dimension of resource allocation is often underexplored. This oversight can lead to GPU fragmentation and suboptimal system performance. In this paper, we propose STS-Fairness, a spatiotemporal sharing fairness scheduler. STS-Fairness partitions each GPU into multiple isolated slots under a novel spatiotemporal fairness constraint and allocates jobs using a round-based allocation mechanism. We guarantee that STS-Fairness achieves overall performance optimality while satisfying spatiotemporal fairness constraints. The scheduling problem is formulated as an integer nonlinear program (INLP) that is solved to optimality in polynomial time via dynamic programming. We deployed the STS-Fairness framework on both physical and simulated heterogeneous clusters and conducted large-scale experiments. These results demonstrate that STS-Fairness reduces average JCT by$1.2 \times$, shortens makespan by$1.24 \times$, and increases throughput by$\mathbf{1. 2 5} \times$compared to state-of-the-art (SoTA) schedulers. Chunhong Du, Mengyu Shi, Shanjiang Tang, Jianhang Tang, Ce Yu, Jian Xiao 0001, Chao Sun 0008, Bin Yang 0043 |
ICPADS | 5 |
| 2025 | MixLoRA: An Efficient Multi-Tenant Framework for Concurrently Serving Diverse LoRA Models in Large Language ModelsabstractThe rapid advancement of large language models (LLMs) has driven the widespread adoption of Low-Rank Adaptation (LoRA) for efficient fine-tuning. However, existing multi-LoRA inference systems, such as Punica, face significant challenges in GPU utilization when handling concurrent requests with diverse ranks, leading to suboptimal throughput and increased latency. To address these limitations, we propose MixLoRA, a novel CUDA stream-based framework that enhances GPU efficiency by enabling parallel execution of mixed-rank LoRA requests. MixLoRA introduces a multi-worker scheduling architecture that selects the optimal number of workers for concurrent execution based on the model size, inference request characteristics, and device information. Each worker manages requests of a fixed rank, eliminating the need for rank-aligned batching and maximizing resource utilization through independent CUDA streams. Experimental results demonstrate that MixLoRA achieves up to 1.2× higher throughput compared to state-of-the-art solutions, offering a scalable and efficient approach for multi-tenant LLM inference services. Ronghuai Chen, Ce Yu, Hao Fu 0021, Xiaoteng Hu, Bin Yang 0043 |
ICPP | 2 |
| 2025 | Solving online resource-constrained scheduling for follow-up observation in astronomy: A reinforcement learning approach
Ce Yu, Chao Sun 0008, Jizeng Wei, Junhan Ju, Shanjiang Tang |
Future Gener. Comput. Syst. | 2 |
| 2025 | Coevolutionary genetic programming for large-scale dynamic multi-aircraft task allocationabstractMulti-aircraft task allocation (MATA) plays a vital role in improving mission efficiency under dynamic conditions. This paper proposes a novel coevolutionary genetic programming (CoGP) framework that automatically designs high-performance reactive heuristics for dynamic MATA problems. Unlike conventional single-tree genetic programming (GP) methods, CoGP jointly develops two interacting populations, i.e., task prioritizing heuristics and aircraft selection heuristics, to explicitly model the coupling between these two interdependent decision phases. A comprehensive terminal set is constructed to represent the dynamic states of aircraft and tasks, whereas a low-level heuristic template translates developed trees into executable allocation strategies. Extensive experiments on public benchmark instances simulating post-disaster emergency delivery demonstrate that CoGP achieves superior performance compared with state-of-the-art GP and heuristic methods, exhibiting strong adaptability, scalability, and real-time responsiveness in complex and dynamic rescue environments. Ce Yu, Xianbin Cao 0001, Wenbo Du 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2025 | Task Scheduling in Geo-Distributed Computing: A SurveyabstractGeo-distributed computing, a paradigm that assigns computational tasks to globally distributed nodes, has emerged as a promising approach in cloud computing, edge computing, cloud-edge computing, and supercomputer computing (SC). It enables low-latency services, ensures data locality, and handles large-scale applications. As global computing capacity and task demands increase rapidly, scheduling tasks for efficient execution in geo-distributed computing systems has become an increasingly critical research challenge. It arises from the inherent characteristics of geographic distribution, including heterogeneous network conditions, region-specific resource pricing, and varying computational capabilities across locations. Researchers have developed diverse task scheduling methods tailored to geo-distributed scenarios, aiming to achieve objectives such as performance enhancement, fairness assurance, and fault-tolerance improvement. This survey provides a comprehensive and systematic review of task scheduling techniques across four major distributed computing environments, with an in-depth analysis of these approaches based on their core scheduling objectives. Through our analysis, we identify key research challenges and outline promising directions for advancing task scheduling in geo-distributed computing. Yujian Wu, Shanjiang Tang, Ce Yu, Bin Yang 0043, Chao Sun 0008, Jian Xiao 0001, Hutong Wu, Jinghua Feng |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2024 | Accuracy-Efficiency Optimization for Multi-Stage Small Object Detection in Surveillance Video with Collaborative Frame SamplingabstractIn video analytics, accuracy and efficiency are two important metrics and there tend to be a tradeoff between each other. In this paper, we consider accuracy-efficiency optimization for small object detection in surveillance video, which is important and has been widely used in many scenarios such as license plate detection in the traffic domain. Given that small objects tend to being attached to big objects, multi-stage object detection is supposed to be an effective approach to achieve high accuracy for small objects by detecting big objects first and then small objects within the ROIs (Region of interests) of big objects. However, existing studies considered the accuracy-efficiency optimization for small object detection only within the single-stage scenario by changing the frame resolution or sampling rate configuration of video data, which are not suitable for multi-stage detection given that its accuracy-efficiency result is determined by the results of all stages jointly. In this paper, we propose an Adaptive and Collaborative frame Sampling approach named ACS for accuracy-efficiency optimization in the multi-stage small object detection. To improve the efficiency significantly while guaranteeing a given accuracy threshold, ACS dynamically adjusts the sampling rates of all stages collaboratively and periodically using the Karush-Kuhn-Tucker (KKT) condition based on the Lagrangian multiplier method. Additionally, we introduce a tuning knob to allow users to flexibly balance accuracy and efficiency, while ensuring a given accuracy threshold λ. Extensive experiments demonstrate the effectiveness of our approach in improving detection efficiency while guaranteeing diverse accuracy requirements. Chunhong Du, Shanjiang Tang, Song Meng, Jiekai Gou, Ce Yu, Yusen Li, Hao Fu 0021 |
CLUSTER | 5 |
| 2024 | A large-scale heterogeneous computing framework for non-uniform sampling two-dimensional convolution applications
Ce Yu, Jian Xiao 0001, Hao Fu 0021, Bo Kang |
CCF Trans. High Perform. Comput. | 2 |
| 2024 | Fast and accurate novelty detection for large surveillance video
Shanjiang Tang, Ce Yu, Chao Sun 0008, Yusen Li, Jian Xiao 0001 |
CCF Trans. High Perform. Comput. | 3 |
| 2024 | Fairness-Efficiency Scheduling for Pay-as-You-Go Shared Caching Systems With Long-Term Fairness GuaranteesabstractPay-as-you-go caching systems are now widely used as storage services in cloud computing. However, users’ data caching requirements not only change over time, but also they are affected by workload characteristics, making it difficult to always ensure high efficient use of cache resources. Cache resource sharing is an effective way to improve the efficiency of cache usage. To incentivize users to share caches, it is essential to ensure long-term fairness among multiple users. However, traditional resource allocation strategies canonly guarantee memory less fairness among users,which is not applicable to long-term cache sharing systems. In this paper, we propose a fair allocation policy named FairCache for Pay-as-you-go cache resources. First, FairCache can satisfy four desirable properties of resource allocation: sharing incentive, pay-as-you-usefairness, strategy proofness, and pare to efficiency. Second, FairCache is an efficiency-fairness resource allocation policy based on the efficiency knob θ. The strategy keeps sensitive to the constantly changing cache demands of multiple users within the system by adjusting the efficiency knob θ, thus ensuring long-term multi-user fairness while maximizing the efficiency of cache usage. In addition, FairCache also has an anti-cheating mechanism to avoid possible free-rider problems when multiple users cache access files. Finally, this paper implements the FairCache policy in Alluxio. The experimental results show that FairCache is a lightweight scheduler and that it can maximize the efficiency usage of cache resources while ensuring the long-term for multiple users in the pay-as-you-go Cache systems fairness. Shanjiang Tang, Zhongyu Zhou, Jiekai Gou, Ce Yu, Yusen Li, Hao Fu 0021, Chao Sun 0008, Jian Xiao 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2024 | EasyRP-R-CNN: a fast cyclone detection model
Xiaoxian Tian, Chongke Bi, Ce Yu |
Vis. Comput. | 4 |
| 2023 | A Survey on Spark Ecosystem: Big Data Processing Infrastructure, Machine Learning, and Applications (Extended abstract)abstractWith the explosive increase of big data in industry and academic fields, it is important to apply large-scale data processing systems to analyze Big Data. Arguably, Spark is the state-of-the-art in large-scale data computing systems nowadays, due to its good properties including generality, fault tolerance, high performance of in-memory data processing, and scalability. Spark adopts a flexible Resident Distributed Dataset (RDD) programming model with a set of provided transformation and action operators whose operating functions can be customized by users according to their applications. It is originally positioned as a fast and general data processing system. A large body of research efforts have been made to make it more efficient (faster) and general by considering various circumstances since its introduction. In this survey, we aim to have a thorough review of various kinds of optimization techniques on the generality and performance improvement of Spark. We introduce various data management and processing systems, machine learning algorithms and applications supported by Spark. Additionally, we make a discussion on the open issues and challenges for large-scale in-memory data processing with Spark. Shanjiang Tang, Bingsheng He, Ce Yu, Yusen Li, Kun Li 0027 |
ICDE | 3 |
| 2023 | Visual Analytics of Air Pollutant Propagation Path and Pollution SourceabstractRecently, controlling air pollution has become increasingly significant due to its impact on our health and daily lives. To prevent and control pollution, it is crucial to trace its source. Many researches have been developed for tracing the source of pollution. However, traditional methods using large-scale simulations need a large number of computation resources and time-consuming. In addition, traditional traceability algorithms do not consider topographic factors, which can cause a certain amount of errors. To resolve above problems, an interactive visual analytics system for pollutant traceability is proposed. In our method, instead of three-dimensional field data, only two-dimensional grid data is enough to track pollution sources in real time. Furthermore, our method can further improve precision through considering topographic factors, which are usually ignored by existing methods. Finally, the possible pollution sources are also identified in our method. This is achieved through analysis of changes in pollutant concentration and the distribution of man-made emission sources. In order to verify the effectiveness of this method, we propose a series of application examples to comprehensively analyze the sources of pollutants. Yan Hao, Chongke Bi, Lu Yang 0007, Xiaobin Qiu, Ce Yu |
VINCI | 6 |
| 2023 | HEGrid: A high efficient multi-channel radio astronomical data gridding framework in heterogeneous computing environments
Ce Yu, Jian Xiao 0001, Shanjiang Tang, Min Long 0001 |
Future Gener. Comput. Syst. | 2 |
| 2022 | EasyNUSC: An Efficient Heterogeneous Computing Framework for Non-uniform Sampling Two-Dimensional Convolution Applications
Ce Yu, Jian Xiao 0001, Hao Fu 0021, Shanjiang Tang, Bo Kang |
ICA3PP | 2 |
| 2022 | Long-Term Fairness Scheduler for Pay-as-You-Use Cache Sharing Systems
Zhongyu Zhou, Shanjiang Tang, Hao Fu 0021, Wanqing Chang, Ce Yu, Chao Sun 0008, Yusen Li, Jian Xiao 0001 |
ICA3PP | 5 |
| 2022 | A method for efficient radio astronomical data gridding on multi-core vector processor
Ce Yu, Jian Xiao 0001, Shanjiang Tang, Hao Fu 0021, Bo Kang, Chenzhou Cui |
Parallel Comput. | 2 |
| 2022 | Fairness-Efficiency Scheduling for Cloud Computing With Soft Fairness GuaranteesabstractFairness and efficiency are two important metrics for users in modern data center computing system. Due to the heterogeneous resource demands of CPU, memory, and network I/O for users’ tasks, it cannot achieve the strict 100 percent fairness and the maximum efficiency at the same time. Existing fairness-efficiency schedulers (e.g., Tetris) can balance such a tradeoff elastically by relaxing fairness constraint for improved efficiency using the knob. However, their approaches areunawareof fairness degradation under different knob configurations, which makes several drawbacks. First, it cannot tell how muchrelaxedfairness can be guaranteed given a knob value. Second, it fails to meet several essential properties such as sharing incentive. To address these issues, we propose a new fairness-efficiency scheduler,QKnober, to balance the fairness and efficiency elastically and flexibly using a tunable fairness knob. QKnober is afairness-sensitivescheduler that can maximize the system efficiency while guaranteeing the$\theta$-soft fairness by modeling the whole allocation as a combination offairness-orientedallocation andefficiency-orientedallocation. Moreover, QKnober satisfies fairness properties of sharing incentive, envy-freeness and pareto efficiency given a proper knob value. We have implemented QKnober in YARN and evaluated it using both testbed and simulated experiments. The results show that QKnober outperforms its alternatives DRF and Tetris by 31.2 and 4.5 percent, respectively. Shanjiang Tang, Ce Yu, Yusen Li |
IEEE Trans. Cloud Comput. | 2 |
| 2022 | A Survey on Spark Ecosystem: Big Data Processing Infrastructure, Machine Learning, and ApplicationsabstractWith the explosive increase of big data in industry and academic fields, it is important to apply large-scale data processing systems to analyze Big Data. Arguably, Spark is the state-of-the-art in large-scale data computing systems nowadays, due to its good properties including generality, fault tolerance, high performance of in-memory data processing, and scalability. Spark adopts a flexible Resident Distributed Dataset (RDD) programming model with a set of provided transformation and action operators whose operating functions can be customized by users according to their applications. It is originally positioned as afastandgeneraldata processing system. A large body of research efforts have been made to make it more efficient (faster) and general by considering various circumstances since its introduction. In this survey, we aim to have a thorough review of various kinds of optimization techniques on the generality and performance improvement of Spark. We introduce Spark programming model and computing system, discuss the pros and cons of Spark, and have an investigation and classification of various solving techniques in the literature. Moreover, we also introduce various data management and processing systems, machine learning algorithms and applications supported by Spark. Finally, we make a discussion on the open issues and challenges for large-scale in-memory data processing with Spark. Shanjiang Tang, Bingsheng He, Ce Yu, Yusen Li, Kun Li 0027 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | DVQShare: An Analytics System for DNN-based Video QueriesabstractApplying deep neural networks (DNNs) to video analytics tasks has drawn attention from both academic and industry communities. However, due to the high computational complexity of DNN models and the explosion of video data, it is challenging to process massive concurrent video queries efficiently and effectively. In this paper, we propose a video analytics system named DVQShare to process DNN-based video queries in a batch mode. The key idea is sharing, including time sharing, spatial sharing, and logical sharing. In principle, sharing across queries can help us reduce the overall amount of frames to be analyzed, which can help us improve the overall performance and reduce the monetary cost. Two modules are designed to process video queries by exploiting the above three sharing opportunities. First, an analysis module is integrated to guide the generation of query processing plans. Within this module, temporal sharing is considered to reuse historical results produced by other queries to remove pending frames that have been analyzed, and spatial sharing is adopted to avoid redundant processing over overlapping video clips. Additionally, we utilize logical sharing to further improve system's overall performance by considering the logical relationship between queries. Second, a query processing engine is devised to execute the query pipeline generated by the analysis module and return the final results. In experiments, we implement a prototype of the DVQShare system based on MXNet, and results show that it can achieve up to 2X performance speedup. Hao Fu 0021, Shanjiang Tang, Ce Yu, Yusen Li, Yanjie Liu |
CCGRID | 3 |
| 2021 | HGP4CNN: an efficient parallelization framework for training convolutional neural networks on modern GPUs
Hao Fu 0021, Shanjiang Tang, Bingsheng He, Ce Yu |
J. Supercomput. | 4 |
| 2020 | Balancing Fairness and Efficiency for Cache Sharing in Semi-external Memory SystemabstractData caching and sharing is an effective approach for achieving high performance to many applications in shared platforms such as the cloud. DRAM and SSD are two popular caching devices widely used by many large-scale data application systems such Hadoop and Spark. Due to the limited size of DRAM as well as the large access latency of SSD (relative to DRAM), there is a trend of integrating DRAM and SSD (called semi-external memory) together for large-scale data caching. Shanjiang Tang, Qifei Chai, Ce Yu, Yusen Li, Chao Sun 0008 |
ICPP | 3 |
| 2020 | Performance optimization of non-equilibrium ionization simulations from MapReduce and GPU acceleration
Jian Xiao 0001, Min Long 0001, Ce Yu |
Parallel Comput. | 3 |
| 2019 | HDF5-Based I/O Optimization for Extragalactic HI Data Pipeline of FAST
Yiming Ji, Ce Yu, Jian Xiao 0001, Shanjiang Tang |
ICA3PP (2) | 2 |
| 2018 | An Efficient Retrieval Method for Astronomical Catalog Time Series Data
Bingyao Li 0001, Ce Yu, Xiaoteng Hu, Jian Xiao 0001, Shanjiang Tang, Lianmeng Li, Bin Ma 0022 |
ICA3PP (1) | 2 |
| 2018 | Correction to: An Efficient Retrieval Method for Astronomical Catalog Time Series Data
Bingyao Li 0001, Ce Yu, Xiaoteng Hu, Jian Xiao 0001, Shanjiang Tang, Lianmeng Li, Bin Ma 0022 |
ICA3PP (1) | 2 |
| 2018 | GpDL: A Spatially Aggregated Data Layout for Long-Term Astronomical Observation Archive
Ce Yu, Chao Sun 0008, Shanjiang Tang, Xiangfei Meng |
ICA3PP (2) | 2 |
| 2018 | A Data-Aware Energy-Saving Storage Management Strategy for On-Site Astronomical Observation at Dome A
Xiaoxiao Lu, Chao Sun 0008, Ce Yu, Ming Che, Zijun Xia, Zhaohui Shang |
ICA3PP (2) | 3 |
| 2018 | HyGrid: A CPU-GPU Hybrid Convolution-Based Gridding Algorithm in Radio Astronomy
Jian Xiao 0001, Ce Yu, Chongke Bi, Yiming Ji |
ICA3PP (1) | 3 |
| 2018 | GLP4NN: A Convergence-invariant and Network-agnostic Light-weight Parallelization Framework for Deep Neural Networks on Modern GPUsabstractIn this paper, we propose a network-agnostic and convergence-invariant light-weight parallelization framework, namely GLP4NN, to accelerate the training of Deep Neural Networks (DNNs) by taking advantage of emerging GPU features, especially concurrent kernel execution. To determine the number of concurrent kernels on the fly, we design an analytical model in the kernel analyzer module and integrate a compact asynchronous resource tracker in the resource tracker module for collecting runtime configurations of kernels with low memory and time overheads. We further develop a runtime scheduler module and a pool-based stream manager for handling GPU work queues in GLP4NN to avoid consuming too many CPU threads or processes while dispatching workloads to GPU devices. In our experiments, we integrate GLP4NN into Caffe to accelerate the batch-based training of four well-known networks on NVIDIA GPUs. Experimental results show GLP4NN is able to achieve a speedup of up to 4X over the original implementation as well as keep the convergence property of networks. Hao Fu 0021, Shanjiang Tang, Bingsheng He, Ce Yu |
ICPP | 4 |
| 2018 | QKnober: A Knob-Based Fairness-Efficiency Scheduler for Cloud Computing with QoS Guarantees
Shanjiang Tang, Ce Yu, Chao Sun 0008, Jian Xiao 0001, Yinglong Li |
ICSOC | 2 |
| 2018 | Long-Term Multi-Resource Fairness for Pay-as-you Use Computing SystemsabstractMany current computing systems such as clouds and supercomputers charge users for their resource usages. A user's demand is often changing over time, indicating that it is difficult to keep the high resource utilization all the time for cost efficiency. Resource sharing is a classical and effective approach for high resource utilization. In view of the heterogeneous resource demands of users' workloads, multi-resource allocation fairness is a must for resource sharing in such pay-as-you-use computing systems. However, we find that, existing multi-resource fair policies such as Dominant Resource Fairness (DRF), implemented in currently popular resource management systems such as Apache YARN [4] and Mesos [23], are not suitable for the pay-as-you-use computing systems. We show that this is because of their memoryless characteristic that can cause the following problems in the pay-as-you-use computing systems: 1). users can get resource benefits by cheating; 2). users might not be able to get the total amount of resources that they are entitled to in terms of their resource contributions. In this paper, we propose a new policy called H-MRF, which generalizes DRF and Asset Fairness with the long-term notion. We show that it can address these problems and is suitable for pay-as-you-use computing systems. We have implemented it into YARN by developing a prototype called MRYARN. Finally, we evaluate H-MRF using both testbed and simulated experiments. The experimental results show that there are about 1.1 ~1.5 sharing benefit degrees and 1.2× ~ 1.8× performance improvement for users with H-MRF, better than existing fair schedulers. Shanjiang Tang, Zhaojie Niu, Bingsheng He, Bu-Sung Lee, Ce Yu |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2017 | KD-Tree and HEALPix-Based Distributed Cone Search Indexing System for Multi-Band Astronomical Catalogs
Ce Yu, Jian Xiao 0001, Xiaoteng Hu, Hao Fu 0021, Kun Li 0001, Yanyan Huang |
ICA3PP | 2 |
| 2017 | Optimized Data Layout for Spatio-temporal Data in Time Domain Astronomy
Ce Yu, Chao Sun 0008, Zhaohui Shang, Jinghua Feng, Jian Xiao 0001 |
ICA3PP | 2 |
| 2017 | Survey on Energy-Saving Technologies for Disk-Based Storage Systems
Ce Yu, Chao Sun 0008, Xiaoxiao Lu, Jian Xiao 0001 |
ICA3PP | 1 |
| 2017 | CMSA: a heterogeneous CPU/GPU computing system for multiple similar RNA/DNA sequence alignmentabstractThe multiple sequence alignment (MSA) is a classic and powerful technique for sequence analysis in bioinformatics. With the rapid growth of biological datasets, MSA parallelization becomes necessary to keep its running time in an acceptable level. Although there are a lot of work on MSA problems, their approaches are either insufficient or contain some implicit assumptions that limit the generality of usage. First, the information of users’ sequences, including the sizes of datasets and the lengths of sequences, can be of arbitrary values and are generally unknown before submitted, which are unfortunately ignored by previous work. Second, the center star strategy is suited for aligning similar sequences. But its first stage, center sequence selection, is highly time-consuming and requires further optimization. Moreover, given the heterogeneous CPU/GPU platform, prior studies consider the MSA parallelization on GPU devices only, making the CPUs idle during the computation. Co-run computation, however, can maximize the utilization of the computing resources by enabling the workload computation on both CPU and GPU simultaneously. This paper presents CMSA, a robust and efficient MSA system for large-scale datasets on the heterogeneous CPU/GPU platform. It performs and optimizes multiple sequence alignment automatically for users’ submitted sequences without any assumptions. CMSA adopts the co-run computation model so that both CPU and GPU devices are fully utilized. Moreover, CMSA proposes an improved center star strategy that reduces the time complexity of its center sequence selection process from O(mn 2) to O(mn). The experimental results show that CMSA achieves an up to 11× speedup and outperforms the state-of-the-art software. CMSA focuses on the multiple similar RNA/DNA sequence alignment and proposes a novel bitmap based algorithm to improve the center star strategy. We can conclude that harvesting the high performance of modern GPU is a promising approach to accelerate multiple sequence alignment. Besides, adopting the co-run computation model can maximize the entire system utilization significantly. The source code is available at https://github.com/wangvsa/CMSA . Chen Wang 0004, Shanjiang Tang, Ce Yu, Quan Zou 0001 |
BMC Bioinform. | 4 |
| 2017 | Online Virtual Machine Placement for Increasing Cloud Provider's RevenueabstractCost savings have become a significant challenge in the management of data centers. In this paper, we show that, besides energy consumption, service level agreement(SLA) violations also severely degrade the cost-efficiency of data centers. We present online VM placement algorithms for increasing cloud provider's revenue. First, First-Fit and Harmonic algorithm are devised for VM placement without considering migrations. Both algorithms get the same performance in the worst-case analysis, and equal to the lower bound of the competitive ratio. However, Harmonic algorithm could create more revenue than First-Fit by more than 10 percent when job arriving rate is greater than 1.0. Second, we formulate an optimization problem of maximizing revenue from VM migration, and prove it as NP-Hard by a reduction from 3-Partition problem. Therefore, we propose two heuristics: Least-Reliable-First (LRF) and Decreased-Density-Greedy (DDG). Experiments demonstrate that DDG yields more revenue than LRF when migration cost is low, yet leads to losses when SLA penalty is low or job arriving rate is high, due to the large number of migrations. Finally, we compare the four algorithms above with algorithms adopted in Openstack using a real trace, and find that the results are consistent with the ones using synthetic data. Laiping Zhao, Zhou Jin 0002, Ce Yu |
IEEE Trans. Serv. Comput. | 4 |
| 2016 | Fast Big Data Analysis in Geo-Distributed CloudabstractAs cloud services grow to span more and more globally distributed datacenters, there is an increasingly need for scheduling algorithms to automatically place tasks across these datacenters. In geo-distributed cloud, the limited WAN bandwidth has become the major bottleneck in fast big data analytics. The scheduling algorithm needs to minimize the global completion time, by jointly optimizing task scheduling and WAN data transfer. In this paper, we model the task scheduling as a community detection problem, with respect to the dependency relations between task, data, and datacenters, and propose a Community Detection-based Scheduling (CDS) algorithm, which is able to minimize the WAN data transfer volume. We utilize the real China-Astronomy-Cloud network to evaluate the proposed algorithms. Experimental results show that we can reduce the total data transfer volume by up to 40.7%, and the global completion time by up to 35.8%, compared with the Hypergraph Partition-based scheduling algorithm and the greedy scheduling algorithm. Laiping Zhao, Chenzhou Cui, Ce Yu |
CLUSTER | 4 |
| 2016 | A new energy-aware task scheduling method for data-intensive applications in the cloud
Congcong Xiong, Ce Yu, Chuanlei Zhang |
J. Netw. Comput. Appl. | 3 |
| 2016 | A general and fast distributed system for large-scale dynamic programming applications
Chen Wang 0004, Ce Yu, Shanjiang Tang, Jian Xiao 0001, Xiangfei Meng |
Parallel Comput. | 2 |
| 2015 | A Multilevel Fault-Tolerance Technique for the DAG Data Driven ModelabstractFault tolerance of hardware failure is a challenging work for parallel programming in massively parallel processing environment. However, traditional rollback-recovery techniques, which an be classified into checkpoint-based and log-based, would introduce extra overhead for recording an overall snapshot of an application. For a specialized programming model, a private recovery technique is valuable and can achieve a better performance.In this paper, a multilevel fault-tolerance technique designed for the DAG data driven model is proposed. It utilized the checkpoint-based fault tolerance technique for system recovery, and timeout to detect and revoery from performance faults. It consists of two kinds of checkpoints: the DAG pattern checkpoint and the intermediate result checkpoint. The DAG pattern checkpoint is designed for tracing the current processing progress of the DAG model, while the intermediate results checkpoint is used to record outputs of compute nodes. Moreover, we also implement this technique in the EasyHPS runtime system. Experimental results show that the check pointing overhead is as low as 2.6%. Hao Fu 0021, Ce Yu |
CCGRID | 2 |
| 2015 | Joint Scheduling of Data and Computation in Geo-Distributed Cloud SystemsabstractRecent trends show that cloud computing is growing to span more and more globally distributed data centers. For geo-distributed data centers, there is an increasing need for scheduling algorithms to place tasks across data centers, by jointly considering data and computation. This scheduling must deal with situations such as wide-area distributed data, data sharing, WAN bandwidth costs and data center capacity limits, while also minimizing completion time. However, this kind of scheduling problems is known to be NP-Hard. In this paper, inspired by real applications in astronomy field, we propose a two-phase scheduling algorithm that addresses these challenges. The mapping phase groups tasks considering the data-sharing relations, and dispatches groups to data centers by way of one-to-one correspondence. The reassigning phase balances the completion time across data centers according to relations between tasks and groups. We utilize the real China-Astronomy-Cloud model and typical applications to evaluate our proposal. Simulations show that our algorithm obtains up to 22% better completion time and effectively reduces the amount of data transfers compared with other similar scheduling algorithms. Lingyan Yin, Laiping Zhao, Chenzhou Cui, Jian Xiao 0001, Ce Yu |
CCGRID | 6 |
| 2015 | A Data Placement Strategy for Data-Intensive Scientific Workflows in CloudabstractWith the arrival of cloud computing and Big Data, many scientific applications with large amount of data can be abstracted as scientific workflows and running on a cloud environment. Distributing these datasets intelligently can decrease data transfers efficiently during the workflow's execution. In this paper, we proposed a 2- stage data placement strategy. In the initial stage, we cluster the datasets based on their correlation, and allocate these clusters onto data centers. Compared with existing works, we have incorporated the data size into correlation calculation, and have proposed a new type of data correlation for the intermediate data named "the first order conduction correlation". Hence the data transmission cost can be measured more reasonable. In the runtime stage, the re-distribution algorithm can adjust data layout according to the changed factors, and the overhead of re-layout itself has also been measured. Compared with previous work, simulation results show that our proposed strategy can effectively reduce the time consumption of data movements during the workflow execution. Congcong Xiong, Ce Yu, Jian Xiao 0001 |
CCGRID | 4 |
| 2015 | A List Scheduling Algorithm for DAG-Based Parallel Computing Models
Hao Fu 0021, Ce Yu |
ICA3PP (2) | 2 |
| 2015 | AQUAdex: A Highly Efficient Indexing and Retrieving Method for Astronomical Big Data of Time Series Images
Zhi Hong, Ce Yu, Ruolei Xia, Jian Xiao 0001, Chenzhou Cui |
ICA3PP (2) | 2 |
| 2015 | Fast 3-Point Correlation Function Approximation on GPU
Chao Sun 0008, Mujin Yang, Ce Yu |
ICA3PP (2) | 3 |
| 2015 | An Energy Efficient Storage System for Astronomical Observation Data on Dome A
Zichao Yuan, Ce Yu, Jian Xiao 0001, Zhaohui Shang |
ICA3PP (4) | 2 |
| 2015 | DPX10: An Efficient X10 Framework for Dynamic Programming ApplicationsabstractX10 language and Asynchronous Partitioned Global Address Space (APGAS) model is an emerging mechanism for programming high-performance computers and commodity clusters. However, little work exists on distributed programming framework for dynamic programming (DP) problems based on X10 and APGAS model. In this paper we present DPX10, an efficient distributed X10 framework for DP applications. DPX10 enables developers to write highly efficient DP programs without much effort. A DPX10 program is specified by a directed acyclic graph (DAG) pattern and a compute method for the vertices. DPX10 provides eight commonly used DAG patterns and a simple API to create custom patterns. The system handles all the tiresome work of implementing parallelization including DAG distribution, vertices scheduling, and vertices communication. Moreover, a new recovery method for distributed arrays is developed to provide transparent fault tolerance. We describe the design of the framework and use four DP applications with up to a billion vertices on 120 cores to demonstrate its simplicity, efficiency, and scalability. Chen Wang 0004, Ce Yu, Xiangfei Meng |
ICPP | 2 |
| 2015 | Accelerating Spectral Calculation through Hybrid GPU-Based ComputingabstractSpectral calculation and analysis have very important practical applications in astrophysics. The main portion of spectral calculation is to solve a large number of one-dimensional numerical integrations at each point of a large three-dimensional parameter space. However, existing widely used solutions still remain in process-level parallelism, which is not competent to tackle numerous compute-intensive small integral tasks. This paper presented a GPU-optimized approach to accelerate the numerical integration in massive spectral calculation. We also proposed a load balance strategy on hybrid multiple CPUs and GPUs architecture via share memory to maximize performance. The approach was prototyped and tested on the Astrophysical Plasma Emission Code (APEC), a commonly used spectral toolset. Comparing with the original serial version and the 24 CPU cores (2.5GHz) parallel version, our implementation on 3 Tesla C2075 GPUs achieves a speed-up of up to 300 and 22 respectively. Jian Xiao 0001, Xingyu Xu 0006, Ce Yu, Jiawan Zhang, Shuinai Zhang |
ICPP | 3 |
| 2014 | A Scalable Real-Time Photometric System for Automatic Astronomical Observations on Dome AabstractDeployed on Dome A of Antarctica, the AST3 astronomical telescopes are required to perform the sky survey in the extreme unmanned environments and process the observed data in real-time to offer the astronomers far back home with the timely observation results. A scalable real-time photometric system is proposed and designed for the automatic astronomical observations of AST3. A GPU-based algorithm for image subtraction photometry is proposed to improve the efficiency of this most time-consuming prodedure in the data processing workflow. To impove the reliability, the system is organized in a HA cluster and equipped with a specially designed daemon. The system has been practically utilized in the astronomical observations with the development of AST3 telescopes. Ce Yu, Lianmeng Li, Jian Xiao 0001, Zhaohui Shang |
CCGRID | 1 |
| 2014 | Probability Based Algorithms for Guaranteeing the Stability of Rechargeable Wireless Sensor Networks
Yiyi Gao, Ce Yu, Jian Xiao 0001, Guiyuan Jiang |
ICA3PP (1) | 2 |
| 2014 | A Platform for Stock Market Simulation with Distributed Agent-Based Modeling
Ce Yu, Hutong Wu, Yuelei Li |
ICA3PP (2) | 2 |
| 2012 | EasyPDP: An Efficient Parallel Dynamic Programming Runtime System for Computational BiologyabstractDynamic programming (DP) is a popular and efficient technique in many scientific applications such as computational biology. Nevertheless, its performance is limited due to the burgeoning volume of scientific data, and parallelism is necessary and crucial to keep the computation time at acceptable levels. The intrinsically strong data dependency of dynamic programming makes it difficult and error-prone for the programmer to write a correct and efficient parallel program. Therefore, this paper builds a runtime system named EasyPDP aiming at parallelizing dynamic programming algorithms on multicore and multiprocessor platforms. Under the concept of software reusability and complexity reduction of parallel programming, a DAG Data Driven Model is proposed, which supports those applications with a strong data interdependence relationship. Based on the model, EasyPDP runtime system is designed and implemented. It automatically handles thread creation, dynamic data task allocation and scheduling, data partitioning, and fault tolerance. Five frequently used DAG patterns from biological dynamic programming algorithms have been put into the DAG pattern library of EasyPDP, so that the programmer can choose to use any of them according to his/her specific application. Besides, an ideal computing distribution model is proposed to discuss the optimal values for the performance tuning arguments of EasyPDP. We evaluate the performance potential and fault tolerance feature of EasyPDP in multicore system. We also compare EasyPDP with other methods such as Block-Cycle Wavefront (BCW). The experimental results illustrate that EasyPDP system is fine and provides an efficient infrastructure for dynamic programming algorithms. Shanjiang Tang, Ce Yu, Bu-Sung Lee, Huabei Wu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2011 | Visual Modeling for Parallel Programming Based on DSLabstractParallel Application Visual Modeling (PAVM) is a system that simplifies the development of parallel applications by providing a graphical user interface for visually modeling and generating corresponding source code framework according to the constructed model. The specification of graphical construction blocks and composition rules are proposed to help construct feasible visual models. Model checker is designed to conduct model verification, and code generator is implemented to translate models to source code framework. PAVM is implemented based on DSL tools. With the help of PAVM, programmers are able to focus on algorithm design and obtain source code framework from graphical models. Ce Yu, Chao Sun 0008 |
CloudCom | 2 |
| 2011 | Fast n-point Correlation Function Approximation with Recursive Convolution for Scalar FieldsabstractIn astrophysics, n-point Correlation Function (n-PCF) is an important tool for computation and analysis, but its algorithmic complex has long been a notorious problem. In this paper we are going to propose two algorithms that are easy to be parallized to compute the n-PCF problem efficiently. The algorithms are based on the definition of recursive convolution for scalar fields (RCSF), and it can be computed using varous fast Fourier Transform (FFT) algorithms in literature. Compared to traditional ways of dealing with this problem, our method is most efficient, for that it can achieve results with point sets as large as 1 billion in less than 1 minute. Moreover, the algorithms are intrinsically appropriate to be used on parallel computing environments such as computer clusters, multi-CPU/GPU super-computers, MapReduce and etc. Better computing environments can deal with better accuracy and time requirements. Ce Yu |
CloudCom | 2 |
| 2009 | A Paralleled Large-Scale Astronomical Cross-Matching Function
Ce Yu, Chenzhou Cui, Liqiang Lv, Jian Xiao 0001 |
ICA3PP | 3 |
| 2007 | EasyPAB: An Extensible IDE Framework for Parallel Applications
Ce Yu, Yanyan Huang, Huabei Wu, Xu Zhen |
APPT | 1 |